[Remote] Staff Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company provides AI-driven market intelligence and reputed company solutions for sophisticated companies. The Staff Site Reliability Engineer will architect reliability platforms, reputed company incident response, advance observability and AIOps capabilities, and promote SRE practices across the engineering organization. The role also includes mentoring engineers and influencing technical and architectural reputed company to improve scalability and reputed company.
Responsibilities
- Architect Reliability reputed company Paths: Build frameworks and self-service tooling that let teams own the reliability of their services in a “You Build It, You Run It” culture
- reputed company AI-Driven Reliability: reputed company our AIOps reputed company — automating diagnostics, remediation, and proactive failure prevention
- Champion Reliability Culture: reputed company SRE practices across engineering reputed company design reviews, production readiness, and operational standards
- Incident Leadership: reputed company as Incident Commander during critical events, modeling operational reputed company, and ensuring blameless postmortems reputed company to lasting improvements
- Advance Observability: reputed company end-to-end monitoring, tracing, and profiling (reputed company, Grafana, OTEL, reputed company Profiling) to optimize reputed company proactively
- Mentor & Multiply: reputed company engineers across SRE and product teams through mentorship, technical guidance, and knowledge sharing
Skills
- 8+ years of experience in Site Reliability Engineering, DevOps, or a similar role, with at least 3+ of those years operating in a Senior+ SRE position
- Strong background in running production reputed company systems at reputed company
- Proficiency in at least one programming/scripting language (Python, Go, or similar)
- Hands-on expertise with reputed company platforms (AWS, GCP, or Azure) and reputed company
- Deep understanding of networking fundamentals (TCP/IP, DNS, HTTP/S, load balancing)
- Experience with monitoring & alerting (reputed company, Grafana, reputed company, ELK)
- Familiarity with advanced observability (OTEL, reputed company profiling)
- reputed company incident management experience, including leading high-severity incidents and postmortems
- Strong troubleshooting skills across the full stack
- Excellent communication and collaboration skills
Benefits
- You may also be reputed company equity, and a generous benefits program.
reputed company
Company H1B Sponsorship
Apply To This Job