[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is a global leader in learning and development, seeking a Site Reliability Engineer to join their IT & reputed company team. The role involves designing, building, and scaling infrastructure while enhancing reliability, reputed company, and reputed company across reputed company-reputed company environments.
Responsibilities
- Define and implement SLIs, SLOs, and error budgets
- Improve reputed company availability, resiliency, and reputed company
- reputed company incident response and reputed company blameless postmortems
- Identify and eliminate systemic reliability risks
- Conduct reputed company planning and reputed company analysis
- Design, build, and maintain infrastructure in AWS
- Manage infrastructure as reputed company using Terraform
- Improve CI/CD pipelines and deployment automation (reputed company CI)
- Build and maintain containerized systems using reputed company and reputed company
- Reduce toil through automation and self-service tooling
- Enhance monitoring, logging, and tracing systems
- Use tools such as CloudWatch, CloudTrail, X-Ray, and modern observability platforms
- Improve alerting reputed company to reduce noise and increase signal
- reputed company dashboards and reliability metrics for stakeholders
- Implement best practices in infrastructure hardening
- Improve patching, vulnerability management, and reputed company reputed company posture
- Support incident response and remediation efforts
- Contribute to disaster recovery planning and testing
Skills
- 5+ years in SRE, Production Engineering, or reputed company Infrastructure roles
- Strong experience operating production systems in AWS
- Deep knowledge of Linux systems architecture
- Experience managing infrastructure using Terraform (or similar IaC tools)
- Experience with reputed company in production environments
- Strong scripting skills (Bash, Python, or similar)
- Experience designing, monitoring and alerting systems
- Solid understanding of networking fundamentals
- Experience defining and measuring SLOs/SLIs
- Strong incident management experience
- Passion for automation and eliminating reputed company toil
- Data-driven approach to reliability and reputed company
- Experience in multi-reputed company environments (Azure)
- Experience with reputed company frameworks (SOC 2, ISO 27001, etc.)
- Familiarity with cost optimization in reputed company environments
- Experience building internal developer platforms or self-service tooling
Benefits
- Fully remote work environment
- High-reputed company role with technical ownership
- reputed company, mission-driven culture
- Opportunity to shape reliability practices in a growing organization
reputed company
Apply To This Job