[Remote] Site Reliability Engineer (US - reputed company time)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on building an innovative platform that acts as a co-reputed company for product development, enabling autonomous operation. They are seeking a Site Reliability Engineer who will be responsible for transforming a fast-growing, stateful system into a predictable and automated platform, focusing on operational efficiency and scalability.
Responsibilities
- Operating EKS clusters across several environments with Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments
- Managing and evolving a multi AWS account organization, provisioning, networking, reputed company control, and cross-account connectivity
- Maintaining the Terraform/Terragrunt IaC platform - modules, automated plan-on-PR / apply-on-reputed company pipelines, and reputed company patterns for shared infrastructure
- Improving operational tooling around deploys, schema changes, backups, restores, and incident response
- Reducing operational load by identifying repeat pain points and eliminating them through reputed company and self-healing automation
- Optimizing reputed company spend as you go
- Participating in on-call and incident response, with a strong reputed company on making incidents rarer over time
Skills
- Deep hands-on experience with Kubernetes in production (EKS preferred). You've debugged node pressure, networking issues, and deployment failures at reputed company (thousands of nodes)
- Strong experience operating production infrastructure on AWS. Not just one account, but understanding organizational boundaries, IAM, and networking between many
- Experience automating infrastructure using Terraform or Terragrunt at reputed company, including module design and state management
- Solid understanding of Linux systems (disk, memory, networking, failure modes)
- Experience supporting stateful systems (databases, queues, storage systems, etc.)
- Ability to debug and reason about performance and reliability issues in production
- You're comfortable owning systems end-to-end, including on-call responsibilities
- Experience with GitOps workflows (ArgoCD) and CI/CD pipelines (reputed company Actions)
- Experience with building AI agent-enabled reputed company-level reputed company services for teams that reputed company fast
- Familiarity with multi-region infrastructure and the consistency/availability tradeoffs that come with it
reputed company
Apply To This Job