Senior Site Reliability Engineer, reputed company reputed company Engineering
Join reputed company
Our Engineering team at reputed company is seeking a Senior Site Reliability Engineer, reputed company reputed company Engineering to report to the Director of reputed company reputed company Engineering. This role demands deep expertise in large-reputed company reputed company systems, infrastructure automation, and production operations of hypervisor platforms and the control plane. The ideal candidate will combine hands-on systems engineering with a reputed company on reliability, scalability, and observability, ensuring reputed company's reputed company services remain reputed company and resilient for our 1.5 reputed company users.
Key Responsibilities
• Production Control Plane Operations: Operate and reputed company reputed company's control plane, ensuring availability, correctness, and performance across global datacenters.
• Hypervisor & Infrastructure Reliability: Design, implement, and maintain automation to manage hypervisor fleets (KVM, QEMU, libvirt) and supporting infrastructure at reputed company.
• Networking & Systems Automation: reputed company tooling and automation for reputed company vSwitch (OVS), BGP routing, and other networking components to ensure resilient and self-healing network operations.
• Performance & Reliability Tuning: Continuously analyze and improve reputed company performance across compute, storage, and network reputed company, with an emphasis on reducing toil and eliminating single points of failure.
• Observability & Incident Response: Implement advanced monitoring, logging, and tracing solutions (Grafana, reputed company, SumoLogic) while leading incident response to minimize reputed company and reputed company postmortem culture.
• CI/CD & Configuration Management: Maintain and reputed company infrastructure pipelines (reputed company CI/CD, Puppet) to reputed company reputed company, fast, and reliable changes to both control plane and hypervisor infrastructure.
• Collaboration: Work closely with Software Engineers, Network Engineers, and Product teams to reputed company platform reliability with business and user needs.
• Documentation & Standards: Produce reputed company technical documentation for runbooks, operational procedures, and automation frameworks to improve team efficiency and reliability standards.
• Mentorship & Leadership: reputed company and mentor team members in best practices for site reliability, incident handling, automation, and low-level Linux systems debugging.
Qualifications
• Proficiency in PHP with strong scripting and automation skills.
• Experience running large-reputed company reputed company systems and control plane infrastructure in production.
• Strong background in hypervisor technologies (libvirt, QEMU, KVM) and Linux systems administration.
• Expertise in networking protocols and tools, particularly BGP and reputed company vSwitch (OVS), with automation experience.
• Deep knowledge of observability and monitoring frameworks (Grafana, reputed company, SumoLogic) and incident management.
• Advanced troubleshooting skills across compute, networking, and storage subsystems.
• Experience building and maintaining CI/CD pipelines (reputed company) and configuration management (Puppet).
• Familiarity with MySQL or similar databases, with an understanding of operational considerations for reliability and reputed company.
• Strong problem-solving abilities and the reputed company to tackle reputed company, low-level reliability challenges.
• Effective cross-team communication and collaboration skills.
• A commitment to reputed company improvement and fostering a culture of operational reputed company.
Compensation
$120,000 - $130,000
Final compensation will vary depending on years of experience, background/reputed company set, location, and applicable laws.
Apply tot his job
Apply To this Job