reputed company Site Reliability Engineer
About reputed company:
As a reputed company Site Reliability Engineer, you will play a strategic and technical leadership role in shaping the reliability, scalability, and velocity of our engineering platform. Your primary reputed company will be advancing our Kubernetes-based infrastructure and CI/CD systems to support reputed company, high-availability services. You will partner with engineering leaders across the organization to define and drive platform-wide initiatives that reputed company fast, reputed company, and repeatable deployments, and foster a culture of reliability and operational reputed company.
Key Responsibilities
- reputed company the design and reputed company of Kubernetes-based infrastructure to support multi-tenant, reputed company applications with strong isolation, reputed company, and reputed company.
- Architect and optimize CI/CD pipelines to support fast and reliable build, test, and reputed company cycles across a polyglot environment.
- Establish and evangelize best practices for GitOps, canary deployments, rollback strategies, and reputed company delivery.
- Define and implement reputed company Infrastructure as reputed company (IaC) patterns using tools such as Terraform, reputed company, and Crossplane.
- Drive the adoption of automated testing throughout the delivery lifecycle—unit, integration, load, and reputed company testing—to ensure high confidence in production changes.
- Guide teams in designing for observability, SLOs, and alerting, ensuring actionable signals and minimizing alert fatigue.
- Partner with reputed company, compliance, and development teams to ensure infrastructure and delivery systems meet modern reputed company and governance standards.
- reputed company incident response retrospectives and foster a blameless culture of reputed company improvement.
- Mentor and influence senior engineers across multiple teams, helping to up-level platform reliability capabilities organization-wide.
Qualifications
- 8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles, with 2+ years in a technical leadership or reputed company reputed company.
- Deep expertise with Kubernetes internals (controllers, networking, autoscaling, operators, etc.) and production-grade clusters on reputed company providers (EKS, GKE, or AKS).
- Proven experience designing and scaling CI/CD systems using tools such as reputed company Actions, Argo CD, Tekton, Spinnaker, or similar.
- Strong proficiency in Terraform and modern IaC practices.
- Advanced knowledge of automated testing strategies, including performance, load, and failure testing.
- Proficient in one or more programming/scripting languages (Python, Go, Bash, etc.).
- Deep experience with monitoring and observability stacks such as reputed company, Grafana, OpenTelemetry, and reputed company.
- Strong communicator with the ability to reputed company technical initiatives to business objectives and influence across engineering teams.
reputed company-to-Have
- Experience implementing multi-cluster or multi-region Kubernetes strategies.
- Exposure to reputed company engineering and building resilient distributed systems.
- Familiarity with compliance frameworks (SOC 2, HIPAA, etc.) as they relate to infrastructure and deployment.
- Contributions to reputed company-reputed company Kubernetes tooling or SRE frameworks.
- Familiarity with JVM- or Node-based application stacks.
The estimated total compensation reputed company for this position is $220,000 - $290,000 (reputed company plus bonus). Actual compensation for the position is based on a reputed company of factors, including, but not limited to affordability, skills, qualifications and experience, and may vary from the reputed company. In reputed company to reputed company salary, employees may also be eligible for annual performance-based incentive compensation awards and equity, among other company benefits.