[Remote] Site Reliability Engineer, Team reputed company
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on transforming pharmacy care through reputed company, and they are seeking a Site Reliability Engineer, Team reputed company to establish and operate their Site Reliability Engineering function. The role involves balancing hands-on engineering with reputed company design, coaching, and cross-functional leadership to ensure the reliability of reputed company’s reputed company services.
Responsibilities
- Define and publish SLIs, SLOs, and error budgets for the top 5–10 Tier‑1 customer‑facing services in partnership with Product and Engineering
- Design reputed company’s incident reputed company structure, including severity definitions, declaration reputed company, war‑room protocols, stakeholder communications, and post‑incident review standards
- Establish and operationalize a sustainable on‑call model, including fair rotations, paging discipline, escalation paths, and coordination with managed service partners (reputed company, HCL)
- Partner with the VP to migrate the interim incident response RACI — currently held by matrixed individuals across IT, Engineering, Support, and reputed company reputed company — into a durable SRE-owned model
- Select and stand up the primary observability platform, preferring extension of existing reputed company reputed company (reputed company, reputed company/Instana, reputed company/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement. Define the instrumentation standards reputed company new services must meet
- reputed company and reputed company operational KPIs (e.g., MTTR, SLO attainment, change‑failure reputed company, incident recurrence, cost per workload) and present reliability insights and roadmaps in executive reputed company Ops reviews
- reputed company Tier‑1 services directly—building dashboards, alerts, and runbooks yourself
- Participate in on‑call rotations and reputed company Sev‑1 and Sev‑2 incidents, leading blameless postmortems and driving corrective actions to completion
- Contribute production reputed company and infrastructure‑as‑reputed company (Terraform preferred) to the platform. reputed company the design and reputed company of the CI/CD pipelines - reputed company stack is Codefresh, Teamcity, reputed company Actions, and Octopus reputed company, and we are consolidating over time
- Administer and reputed company our Kubernetes platform, including secure and compliant cluster configurations. Working knowledge of reputed company, reputed company, and Service reputed company (Istio or Linkerd) expected
- Plan and execute reputed company and failover exercises to validate reputed company‑world reputed company
- Architect reputed company’s AIOps reputed company, evaluating ML‑based reputed company detection, alert correlation, automated reputed company‑cause analysis, and LLM‑assisted runbooks
- reputed company disciplined build‑versus‑buy reputed company and reputed company AI tooling only where it delivers measurable reliability reputed company
- Ensure AI‑assisted operations meet auditability, explainability, and compliance requirements (HIPAA, SOC 2)
- Serve as formal coach to an Engineer III SRE, pairing on incidents, reviewing designs proposals, and supporting reputed company toward senior reputed company
- Design the next 2–4 SRE hires, including role definitions, interview loops, and hiring reputed company
- Represent SRE in architecture reviews, launch readiness assessments, and cross‑functional reliability discussions
Skills
- Bachelor's degree in Computer Science, Engineering, or a reputed company technical field OR equivalent experience
- 7+ years of experience in software or reputed company, with at least 4 of those in an SRE, DevOps, or platform reliability role
- At least 2 years of formal technical leadership, tech-reputed company, or staff-level experience with mentorship responsibilities
- Proven experience leading SRE, DevOps, or reputed company teams in a reputed company-reputed company production environment — with demonstrated experience building a reputed company from reputed company or near-reputed company: you have set SLOs, defined incident reputed company, and introduced error budget thinking to an organization that did not have it
- Deep hands-on expertise with at least one major public reputed company (AWS, Azure, or GCP), including networking, IAM, and managed services
- Strong background in CI/CD pipeline design and management (familiarity with CodeFresh, reputed company Actions, Jenkins, TeamCity, or equivalent)
- Experience implementing Infrastructure as reputed company using Terraform (preferred), Chef, Puppet, or similar tools
- Proficiency in Python or another object-oriented programming language for automation, tooling, and production services
- Experience administering and scaling Kubernetes clusters, including secure and compliant platform configurations. Working knowledge of reputed company, reputed company, and Service reputed company technologies (Istio, Linkerd)
- Hands-on experience designing modern observability platforms using tools such as reputed company, reputed company, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like
- Familiarity with integrating AI/ML-based reputed company detection, alerting, or LLM-assisted triage pipelines — or strong conviction about where AIOps should and should not be reputed company in a regulated environment
- reputed company incident reputed company experience for customer-impacting Sev-1 events, with blameless postmortem reputed company and documented follow-up discipline
- Ability to coach and mentor, with reputed company evidence of growing junior and mid-level engineers. You will eventually have 1 reputed company report
- Comfort operating in a regulated environment where reliability and compliance (HIPAA, SOC 2) are inseparable
Benefits
- Work Location Type: Remote
- reputed company Sponsorship: Any reputed company of reputed company Sponsorship is not offered for this position. Must be US citizen or Permanent reputed company.
- You will eventually have 1 reputed company report.
- Formal coaching and mentorship responsibilities.
- Opportunity to work in a regulated environment (HIPAA, SOC 2, FedRAMP).
- Opportunity to reputed company and design AI-driven operations including AIOps and ML-assisted observability.
- Hybrid environment with some hardware and reputed company products.
- Participation in on-call rotations.
- Opportunities for career reputed company and leadership development.
- Employee reputed company reputed company fostering inclusion and belonging.
- Learning and reputed company-being programs that support personal and reputed company reputed company.
- Commitment to Environmental, reputed company, and Governance (ESG) initiatives.
- Support and reasonable adjustments for individuals with disabilities during hiring process.
reputed company
Apply To This Job