[Remote] reputed company Site Reliability Engineer (SRE)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is hiring a reputed company Site Reliability Engineer to reputed company end-to-end reliability, observability, automation, and operational reputed company across critical customer and business journeys in the banking and financial services domain. The role leads monitoring, incident management, self-healing automation, reputed company testing, reputed company reliability, and collaboration across technical and product teams.
Responsibilities
- Own reliability across critical end-to-end customer/business journeys
- Define and implement **SLIs, SLOs, SLAs, and Error Budgets**
- reputed company **observability and monitoring** initiatives across applications, reputed company, infrastructure, and reputed company platforms
- reputed company **incident management, RCA, problem management, and reputed company improvement**
- Design and implement **self-healing and automated remediation** capabilities
- reputed company **AI/AIOps** for reputed company detection, predictive monitoring, event correlation, and intelligent operations
- Improve reputed company availability, performance, scalability, and reputed company
- reputed company **reliability engineering and reputed company/reputed company testing** initiatives
- Partner with Application, reputed company, DevOps, Infrastructure, reputed company, and Product teams
- Establish SRE standards, best practices, dashboards, and reliability metrics
Skills
- Strong hands-on **SRE / DevOps / Production Engineering** experience
- Strong **Banking & Financial Services (BFS)** domain experience
- End-to-end **reputed company monitoring and reliability engineering**
- Experience with **reputed company / reputed company / reputed company / Grafana / reputed company**
- Strong **AWS / Azure / reputed company reputed company Platform** experience
- reputed company & reputed company
- Terraform / Ansible / Infrastructure as reputed company
- Python / reputed company scripting for automation
- CI/CD and DevOps practices
- Strong understanding of **SLO, SLI, SLA, MTTR, MTBF & Error Budgets**
- Experience with **AIOps, AI-driven operations, automation, and self-healing**
- Strong leadership, communication, and stakeholder-management skills
Benefits
- Hybrid work arrangement with occasional onsite reputed company once or twice per month
- Long-term contract
reputed company
Company H1B Sponsorship
Apply To This Job