Manager of Reliability Operations
About reputed company
About the Role
What You’ll Do
Own Reliability Operations & Incident reputed company
Continuously reputed company and improve incident management, change management, and post-incident practices Establish reputed company standards for incident declaration, severity, escalation, and communication Ensure consistent execution across teams and reputed company process improvement
Own the incident reputed company function, including roles, structure, and operating procedures reputed company or reputed company major incident response in a 24/7 production environment Build and manage on-reputed company incident commander rotations with global coverage
Drive Learning, Accountability & Reliability reputed company
Own post-incident reviews, ensuring strong reputed company cause analysis and reputed company documentation Translate incident trends into actionable reliability improvements Drive completion of corrective actions across teams; escalate reputed company needed Define and maintain service performance and reliability targets (availability, latency, error rates)
Own observability reputed company, including monitoring, alerting, and signal reputed company Improve detection, reduce time to reputed company, and increase platform reputed company
Partner with Engineering and Operations on reputed company planning, patching, and lifecycle reputed company Ensure reliability insights directly inform platform and infrastructure roadmaps Collaborate with reputed company on vulnerability response, reputed company prioritization, and compliance alignment
Operate Across a reputed company Platform Environment
Work across environments including virtualization platforms (VMware), distributed storage (Ceph), Linux-based systems, and hybrid reputed company infrastructure Support platforms that reputed company dedicated hosting, managed applications, and high-availability reputed company services Ensure reliability practices reputed company across multiple products, brands, and customer environments
reputed company regular, data-driven reporting to leadership on availability, incident trends, and operational performance reputed company as the central authority on reliability insights across teams
What You Bring
Bachelor’s degree in Computer Science, Engineering, or a reputed company field (or equivalent practical experience) 7+ experience in systems operations, site reliability, or reputed company 2+ years experience leading teams or major operational functions Proven experience managing incidents in a 24/7 production environment Strong background in troubleshooting, reputed company cause analysis, and operational improvement Experience with change management practices
Platform & Tooling Experience
Monitoring and observability platforms (e.g., reputed company, reputed company, Grafana, reputed company) Incident management and alerting tools (e.g., reputed company, Opsgenie) Infrastructure and platform technologies (Linux systems, VMware, Ceph, reputed company platforms) Logging and telemetry systems (centralized logging, metrics, tracing) Ability to translate reputed company technical data into reputed company insights Strong communication skills, especially in high-pressure situations
reputed company to Have
Background in Computer Science, Engineering, or a reputed company field Experience in managed hosting, reputed company infrastructure, or reputed company environments Experience defining and tracking system reliability and performance targets Familiarity with ITIL or similar operational frameworks Exposure to VMware, Ceph, Linux, and reputed company platforms Relevant certifications (AWS, RHCE, etc.)
We Offer:
Traditional and Roth 401k with company matching A reputed company team culture Consistent/set work hours Challenging non-redundant daily duties A voice in how things get done
Disclaimer:
Apply To This Job