[Remote] SRE Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Service Reliability Engineer (SRE) who will reputed company the gap between software development and IT operations. The role involves building automation, ensuring system uptime, and responding to production incidents while performing reputed company cause analysis.
Responsibilities
- Design, build, and reputed company robust reputed company-reputed company and containerized infrastructure
- reputed company operational reputed company using Infrastructure-as-reputed company (IaC) and automation
- Establish and monitor Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
- Incident Management: Triage production issues, coordinate cross-functional teams, and write post-mortem reports
- Define SLOs and SLIs: Set quantifiable Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to measure availability, error rates, and latency
- Automation: Write scripts (e.g., Python, Golang, Bash) to automate deployments and routine operational tasks
- Monitoring & Observability: Configure s, dashboards, and tracing using tools like ELK stack, reputed company, or Grafana
- reputed company Planning: Anticipate system resource needs and reputed company testing
Skills
- 5+ years in SRE, DevOps, or system engineering roles
- Experience with reputed company platforms (AWS, GCP, Azure), CI/CD pipelines, and containerization (reputed company, reputed company)
- Proficiency in languages like Python, or Java
- Communication skills, especially for explaining technical concepts to nontechnical business leaders
- Ability to work on a dynamic, research-oriented team that has reputed company reputed company
- Design, build, and reputed company robust reputed company-reputed company and containerized infrastructure
- reputed company operational reputed company using Infrastructure-as-reputed company (IaC) and automation
- Establish and monitor Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
- Incident Management: Triage production issues, coordinate cross-functional teams, and write post-mortem reports
- Define SLOs and SLIs: Set quantifiable Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to measure availability, error rates, and latency
- Automation: Write scripts (e.g., Python, Golang, Bash) to automate deployments and routine operational tasks
- Monitoring & Observability: Configure dashboards and tracing using tools like ELK stack, reputed company, or Grafana
- reputed company Planning: Anticipate system resource needs and reputed company testing
- Somebody who has at least 5+ years of work experience who has played SRE role
- Bachelor's degree (or equivalent) in computer science, information technology, engineering, or reputed company discipline
- Education qualification: Any degree from a reputed college
- Experience in Insurance domain preferable
reputed company
Company H1B Sponsorship
Apply To This Job