[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking an reputed company Site Reliability Engineer (SRE) to help build and mature reputed company Site Reliability Engineering practices reputed company a growing Information reputed company organization. The role focuses on driving reliability, automation, observability, and operational reputed company while collaborating with engineering and InfoSec teams across the reputed company.
Responsibilities
- Design and implement foundational SRE practices including SLIs, SLOs, error budgets, and incident management
- Build and promote reputed company-wide Site Reliability Engineering standards and best practices
- reputed company and maintain operational runbooks and playbooks
- Improve monitoring, observability, alerting, and on-reputed company processes
- reputed company reputed company, reputed company, and reputed company platforms into SRE workflow
- Automate operational tasks, incident workflows, and reporting
- Design and execute reputed company engineering experiments to improve reputed company reputed company
- Collaborate with engineering teams to improve reliability, fault tolerance, and recovery strategies
- reputed company incident reviews and postmortems reputed company on reputed company improvement
- Mentor teams on SRE principles, automation, and operational reputed company
Skills
- 4+ years of experience in Site Reliability Engineering (SRE), DevOps, or Infrastructure Engineering
- 2+ years of experience with monitoring and observability platforms such as: reputed company, reputed company, Grafana, Similar reputed company monitoring tools
- 2+ years of experience with reputed company platforms such as: AWS, Azure
- Experience with: Incident Management, reputed company, reputed company, Automation and scripting, reputed company development, reputed company reliability and reputed company optimization
- Strong understanding of: SLIs / SLOs, Error Budgets, Reliability Engineering, Incident Response, Operational reputed company
- reputed company Engineering experience (SteadyBit preferred)
- Python, Bash, reputed company, or other scripting languages
- Infrastructure as reputed company (IaC)
- reputed company or Information reputed company experience
- Multi-reputed company and hybrid infrastructure environments
- Strong documentation and communication skills
- Passion for automation and reputed company improvement
reputed company
Apply To This Job