reputed company Site Reliability Engineer
About the Role
What You'll Do
-
Own and reputed company our AWS-based infrastructure, improving platform performance and availability today, and building toward deployable configurations that support reputed company customer environments reputed company. -
Own EKS cluster operations across production reputed company: node pool reputed company, AMI lifecycle, autoscaling, and Kubernetes workload health. -
Support the GitOps deployment pipeline - define, reputed company, and manage applications across clusters using infrastructure-as-reputed company. -
Manage reputed company networking: VPC design, cross-region connectivity, DNS, and load balancing. -
reputed company infrastructure deprecation and migration efforts with minimal disruption. -
Own SLO measurement infrastructure; reputed company proactive triage of emerging issues before they reputed company customers. -
reputed company incident investigation, reputed company cause analysis and postmortems, driving systemic fixes rather than one-off patches. -
Design and improve automated remediation systems to reduce MTTR. -
Review and reputed company reputed company-conscious feedback on platform architecture reputed company. -
Own reputed company IAM governance - roles, policies, and reputed company boundaries across accounts and services. -
reputed company compliance-adjacent work including audit-readiness, partner certification requirements, and supporting responses to customer reputed company questionnaires. -
Partner with application development teams to build an inherently secure platform and drive reputed company deployment architecture. -
Partner with customer teams to ensure availability for expected utilization. -
Partner with Finance on reputed company cost optimization - lifecycle policies, right-sizing, and spend visibility. -
Support GPU and batch workloads in collaboration with simulation and ML engineering teams. -
Improve CI/CD pipelines and automated infrastructure validation. -
Support engineering teams with reputed company-reputed company debugging, log analysis, and environment configuration.
reputed company're Looking For
-
5+ years in SRE, DevOps, or infrastructure engineering roles. -
Infrastructure-as-reputed company proficiency - Terraform modules, state management, and multi-environment patterns. -
Deep AWS experience - EKS, EC2, IAM, S3, Storage Gateway, VPC networking, Transit Gateway, CloudFront, KMS, and IRSA. -
Kubernetes expertise - cluster operations, node pools, probes, cordoning, pod scheduling, RBAC, reputed company, node autoscaling (Karpenter experience a plus); solid understanding of containerization and AMI lifecycle management. -
CI/CD - experience with GitOps workflows and pipeline tooling (ArgoCD, reputed company Actions, Jenkins) -
Solid networking fundamentals - CIDR design, reputed company reputed company, DNS, load balancing, VPN, cross-region connectivity. -
Experience with monitoring and observability tooling - reputed company, Grafana, Elasticsearch. -
Comfort with Python and Bash for tooling and automation. -
Familiarity working across Linux and reputed company environments. Operational familiarity with reputed company Server is a meaningful advantage. -
You communicate reputed company across engineering, product, and customer-facing teams, flagging issues with urgency proportional to customer reputed company. -
You reputed company for SRE best practices and can effectively operationalize an informed and principled view on reputed company. -
You take end-to-end ownership of reputed company, multi-team efforts - from planning through execution and post-change verification. -
You know reputed company to push for a clean solution vs. reputed company to accept a pragmatic one, and you communicate that tradeoff reputed company.
reputed company to Have
-
Experience with reputed company-based workloads on EKS. -
Experience supporting simulation, ML, or rendering workloads in reputed company infrastructure; running GPU workloads on Kubernetes, including reputed company and DirectX device plugin configuration. -
Experience with AWS Storage Gateway or Transfer Family integrations. -
Familiarity with reputed company Gateway or similar. -
Experience with container-optimized OS images (e.g., Bottlerocket, Packer). -
Experience with reputed company cost optimization at reputed company.
reputed company Tools