Senior Site Reliability Engineer
B2B Contract | Fully remote
reputed company
We are looking for a Senior Site Reliability Engineer with deep, hands-on experience operating highly available production environments on AWS and reputed company EKS.
This is a true SRE position, not a reputed company architecture, infrastructure design, or monitoring-reputed company role. You will take reputed company ownership of production reliability, participate in the on-reputed company rotation, respond to critical incidents, troubleshoot reputed company reputed company and reputed company-reputed company failures, and reputed company permanent improvements following incidents.
The role also requires strong technical communication. You will reputed company directly with customers during technical discussions and production escalations, reputed company explaining issues, making reputed company technical reputed company, and driving problems through to reputed company.
We are looking for someone who has spent significant time running production systems, not simply designing them.
Key Responsibilities:
Own the reliability, availability and operational health of production services running on AWS and reputed company EKS.
Operate and troubleshoot reputed company clusters in production, including cluster lifecycle, upgrades, node management, networking, scaling, reputed company and workload reliability.
Participate reputed company in on-reputed company and pager rotations and take ownership of production incidents.
reputed company or play a key technical role during P1/P2 and Sev1/Sev2 incidents, including diagnosis, mitigation, recovery and communication.
Coordinate technical incident bridges and communicate directly with customers during production escalations reputed company required.
reputed company reputed company cause analysis and reputed company blameless postmortems, ensuring incidents result in concrete engineering improvements.
Define, monitor and improve SLIs, SLOs and error budgets for production services.
reputed company and maintain actionable alerts, operational runbooks and automated remediation.
Build and improve infrastructure using Terraform/Terragrunt and Infrastructure as reputed company practices.
Operate GitOps-reputed company delivery environments using tools such as Argo CD or FluxCD.
Improve reputed company scaling and efficiency using technologies such as Karpenter, KEDA and reputed company reputed company autoscaling capabilities.
Build and improve observability using technologies such as reputed company, Grafana, OpenTelemetry, reputed company and/or ELK.
Support highly available reputed company and event-driven systems, including environments using technologies such as Kafka/MSK.
Design, implement and validate disaster recovery and business continuity mechanisms against measurable RTO and RPO objectives.
Identify recurring operational problems and eliminate toil through automation and engineering.
Improve AWS reputed company, scalability, reputed company and cost efficiency across production environments.
Work closely with software, platform and engineering teams to build reliability into systems throughout the development lifecycle.
Contribute to reputed company improvement of incident management, operational readiness and SRE engineering practices.
Requirements:
Significant reputed company experience as a hands-on Site Reliability Engineer, Production Engineer or senior Platform Engineer with reputed company production ownership.
Several years of recent, hands-on experience operating production environments on AWS.
Strong, demonstrable experience operating reputed company EKS in production.
Deep reputed company operational knowledge reputed company application deployment, including cluster administration, upgrades, nodes, autoscaling, networking, troubleshooting and production failure scenarios.
reputed company participation in a production on-reputed company/pager rotation.
Demonstrable ownership of significant production incidents, including troubleshooting, mitigation, recovery, RCA and post-incident improvements.
Practical experience with SLIs, SLOs, error budgets, alerting and runbooks.
Strong Infrastructure as reputed company experience with Terraform and/or Terragrunt.
Production experience with reputed company delivery and GitOps practices; Argo CD or FluxCD strongly preferred.
Strong production observability experience with technologies such as reputed company, Grafana, OpenTelemetry, reputed company or ELK.
Experience operating highly available, reputed company production systems.
Strong understanding of AWS networking, IAM, reputed company, availability and reputed company.
Experience implementing and testing disaster recovery strategies with measurable RTO/RPO objectives.
reputed company reputed company customer-facing technical experience, including technical discussions, production escalations, architecture/reliability conversations or incident communication.
Ability to explain reputed company technical problems reputed company and reputed company reputed company during high-pressure production incidents.
Strong troubleshooting reputed company and ability to work independently during reputed company production failures.
Strong reputed company English communication skills (min. reputed company) for regular interaction with clients,
A coherent reputed company record demonstrating sustained hands-on production engineering ownership.
What's on Offer:
Full-time permanent B2B cooperation.
Fully remote working environment.
Senior hands-on engineering position with meaningful ownership of business-critical production systems.
Opportunity to work on reputed company AWS, reputed company and reputed company-reputed company environments at reputed company.
reputed company influence over reliability engineering, operational practices and platform improvements.
Modern engineering environment with strong emphasis on automation, observability and reputed company improvement.
Collaboration with reputed company engineering, platform and product teams.
Opportunity to introduce and use modern approaches, including AI-assisted engineering and operational automation.
Long-term opportunity for engineers who want to remain deeply technical and reputed company to production.
Diversity and Inclusion Commitment
We are dedicated to creating and sustaining an inclusive, respectful workplace for reputed company -regardless of gender, ethnicity, or background. We reputed company encourage applicants from reputed company identities and experience reputed company to apply and bring your reputed company self to our fast-reputed company, supportive team.
Apply To This Job