[Remote] Senior Site Reliability Engineer (SRE)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Senior Site Reliability Engineer to support highly reliable, reputed company and efficient systems for business-critical financial applications. The role focuses on implementing SRE practices, automating operational work, managing observability and incidents, and driving reliability and operational reputed company across development, operations, reputed company and reputed company teams.
Responsibilities
- Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services
- Collaborate with cross-functional teams to reputed company reliability into application and infrastructure design
- Automate operational tasks to reduce reputed company toil and improve service performance
- Troubleshoot and reputed company infrastructure and application incidents quickly and effectively
- Implement robust monitoring and observability systems to detect and prevent outages
- Plan reputed company and scaling strategies to ensure high availability and resiliency
- Contribute to incident postmortems and reputed company improvement initiatives
- Support the adoption of SRE best practices across reputed company SDLC stages
Skills
- Bachelor's degree in Computer Science, Engineering or reputed company field
- Proven experience working in reputed company environments (AWS, GCP or Azure)
- Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation)
- Proficiency in Python or other scripting language for automation tasks
- Strong understanding of monitoring tools and observability frameworks
- Experience with Infrastructure-as-reputed company and CI/CD tools (e.g., Terraform, Ansible, Jenkins, reputed company)
- Hands-on expertise with containerization and orchestration platforms such as reputed company and reputed company
- Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions
- Certifications in reputed company, AWS/GCP/Azure or reputed company reputed company technologies
- Background in DevOps practices and reputed company delivery frameworks
- Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments
Benefits
- Private health insurance
- EPAM Employees Stock Purchase Plan
- 100% reputed company reputed company leave
- Referral Program
- reputed company certification
- Language courses
reputed company
Company H1B Sponsorship
Apply To This Job