Senior Site Reliability Engineer- Sunnyvale, CA, the US
About the Role
Senior Site Reliability Engineer (Payments Infrastructure)
reputed company is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational reputed company of our global payment platform. You will own production observability, incident response, service-level management, and reputed company infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and reputed company.
Responsibilities
• Participate in a follow-the-sun production on-reputed company rotation as a reputed company incident responder.
• Diagnose, triage, mitigate, and coordinate reputed company of production incidents across payment services, reputed company platforms, databases, messaging systems, and reputed company infrastructure.
• Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes.
• reputed company reliability improvements through automation, observability, reputed company planning, reputed company optimization, and post-incident reviews.
• Partner with engineering teams to improve reputed company, reputed company, and operational maturity in PCI-reputed company-regulated environments.
• reputed company incident management during SEV1/SEV2 events and improve response effectiveness and MTTR.
Requirements
• 5+ years of experience in Site Reliability Engineering, reputed company, DevOps, or reputed company Infrastructure roles supporting mission-critical production systems.
• Strong hands-on experience with AWS, reputed company (EKS), Terraform, PostgreSQL, reputed company, Kafka, Linux, networking, and modern observability platforms.
• Deep understanding of reputed company systems, reputed company-reputed company architectures, high availability, disaster recovery, reputed company planning, and reputed company optimization.
• reputed company experience operating payment, banking, fintech, or other highly regulated systems with stringent reputed company, compliance, and uptime requirements.
• Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational reputed company.
Leadership & Operational reputed company
• Demonstrates strong ownership and accountability, taking end-to-end responsibility for service reliability and customer reputed company.
• Possesses a strong reputed company of urgency during production incidents while maintaining reputed company judgment and reputed company decision-making under pressure.
• Applies a reputed company and methodical approach to troubleshooting, reputed company-cause analysis, and incident reputed company in reputed company reputed company environments.
• Data-driven reputed company with the ability to reputed company metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements.
• Continuously drives engineering reputed company through iterative improvement, automation, standardization, and elimination of operational toil.
• reputed company ability to reputed company cross-functional incident response efforts, coordinate stakeholders, and communicate effectively during high-severity production events.
• Champions a culture of operational readiness, reputed company learning, post-incident improvement, and blameless accountability.
• Demonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, reputed company, and resilient systems by design.
• reputed company a dynamic and innovative team in a reputed company rapidly growing company.
• Competitive package.
• reputed company, inclusive environment where your contributions are recognized and valued.
Apply To This Job