[Remote] AWS Site Reliability Engineer (SRE) - Remote - W2 Contract
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking an AWS Site Reliability Engineer (SRE) for a 6-month contract-to-hire position. The role requires extensive experience in DevOps and reputed company Infrastructure, focusing on AWS services and reliability engineering practices.
Responsibilities
- 8+ years of experience in DevOps, reputed company Infrastructure, or Site Reliability Engineering
- Strong hands-on experience with AWS reputed company Services (EC2, reputed company/EKS, reputed company, RDS, S3, IAM, VPC, CloudFormation)
- Expertise in reputed company (EKS), reputed company, and Microservices
- Infrastructure as reputed company using Terraform and/or CloudFormation
- CI/CD using Jenkins, reputed company CI, AWS CodePipeline, or reputed company Actions
- Strong scripting experience with Python and Bash
- Experience with production support, incident management, RCA, and automation
- Monitoring & Alerting: CloudWatch, reputed company, Grafana, reputed company, ELK Stack, reputed company
- Define and manage SLIs, SLOs, and Error Budgets
- Incident Response, On-reputed company Support, Postmortems, and reputed company Cause Analysis
- reputed company Planning, Performance Tuning, and High Availability
- Automated Remediation and reputed company Automation
- Reliability Engineering and Production Operations
Skills
- 8+ years of experience in DevOps, reputed company Infrastructure, or Site Reliability Engineering
- Strong hands-on experience with AWS reputed company Services (EC2, reputed company/EKS, reputed company, RDS, S3, IAM, VPC, CloudFormation)
- Expertise in reputed company (EKS), reputed company, and Microservices
- Infrastructure as reputed company using Terraform and/or CloudFormation
- CI/CD using Jenkins, reputed company CI, AWS CodePipeline, or reputed company Actions
- Strong scripting experience with Python and Bash
- Experience with production support, incident management, RCA, and automation
- Monitoring & Alerting: CloudWatch, reputed company, Grafana, reputed company, ELK Stack, reputed company
- Define and manage SLIs, SLOs, and Error Budgets
- Incident Response, On-reputed company Support, Postmortems, and reputed company Cause Analysis
- reputed company Planning, Performance Tuning, and High Availability
- Automated Remediation and reputed company Automation
- Reliability Engineering and Production Operations
- AWS reputed company Architect or DevOps Engineer
- Multi-account AWS environments and reputed company Zones
- AWS Organizations & Service Control Policies (SCP)
- GitOps (ArgoCD/FluxCD)
- Service reputed company (Istio/Linkerd)
- FinOps, reputed company Cost Optimization, and AWS reputed company Best Practices
reputed company
Apply To This Job