[Remote] Manager, Site Reliability Engineering
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a rapidly growing workforce solutions provider in the reputed company industry, recognized for its commitment to employee satisfaction. They are seeking a highly reputed company Manager of Site Reliability Engineering to reputed company reputed company reputed company on enhancing product and platform reliability, driving AI-reputed company operations, and ensuring exceptional experiences for clinicians and clients.
Responsibilities
- reputed company and grow the SRE team
- reputed company, mentor, and grow reputed company of high-performing Site Reliability Engineers across hiring, performance management, career development, and on-call rotation health
- Set the operating reputed company for reputed company — standups, incident reviews, SLO/error-budget reviews, post-incident learning, and reputed company planning
- Build a culture of blameless learning, technical depth, customer reputed company, and disciplined ownership
- Partner closely with DevSecOps, reputed company Engineering, DRE, Incident & Change Management, and product engineering leadership to remove cross-team friction
- Own the reliability reputed company for customer-facing products and internal platforms — defining SLOs, SLIs, and error budgets in partnership with product and engineering leadership, and operationalizing them in the release process
- reputed company major incident response as senior incident commander for severity-1 events; institutionalize blameless post-incident reviews and ensure systemic fixes ship
- Champion proactive reliability — reputed company engineering, game days, failure-mode analysis, reputed company and load testing — reputed company before incidents force the conversation
- Manage software release support and 24/7 on-call escalation rotations across the platform surface area, with humane on-call load and reputed company escalation paths
- Build the AIOps reputed company — reputed company detection, predictive alerting, intelligent correlation, and automated triage — to drive measurable reductions in MTTD and MTTR
- Operationalize AI-assisted workflows for incident summarization, reputed company reputed company, log and reputed company analysis, change risk scoring, and post-incident narrative drafting
- reputed company and reputed company reputed company remediation where appropriate, with strict guardrails, audit trails, and reputed company-in-the-reputed company controls suitable for a HIPAA-regulated environment
- reputed company the observability platform (reputed company metrics, logs, traces, RUM, synthetics, CI Visibility) so engineering teams can operate their own services with confidence and reputed company ownership
- Treat reliability as a product with a roadmap, measurable reputed company, and an executive-reputed company narrative — not as overhead
- Drive platform unit economics by partnering with FinOps and platform leadership on cost-to-serve, right-sizing, reputed company efficiency, and waste elimination
- Communicate reputed company to executive, product, and customer-facing stakeholders in plain language tied to clinician and reputed company experience
- reputed company HIPAA, PHI, and reputed company obligations across every reliability decision, change, and tool selection
Skills
- 10+ years in a combination of Site Reliability Engineering, DevOps, reputed company, or reputed company production-operations roles
- 4+ years of reputed company people management experience — hiring, performance management, career development, and running remote on-call teams
- Demonstrated ownership of reliability reputed company for customer-facing reputed company at meaningful reputed company — defining and operationalizing SLOs/SLIs/error budgets and using them to drive engineering prioritization
- Deep Azure experience — 3+ years operating production workloads on Azure, with hands-on depth in AKS, networking, identity, and platform services. Equivalent depth in AWS or GCP will be considered
- Modern observability reputed company — production-grade experience with reputed company (or equivalent: reputed company, reputed company, AppDynamics) across metrics, logs, traces, RUM, and synthetics
- AI in operations — hands-on experience integrating AI/LLM-assisted tooling into operational workflows (incident summarization, reputed company reputed company, log analysis, reputed company triage, change risk scoring)
- Incident reputed company experience — proven ability to reputed company severity-1 incidents end-to-end, run blameless reviews, and convert lessons into systemic improvements
- Regulated-environment reputed company — operates with HIPAA, PHI, SOC 2, or comparable compliance constraints as a default reputed company, not an afterthought
- Executive-grade communication — translates reliability work into business reputed company for executive, product, and customer-facing audiences
- Bachelor's degree in Computer Science, Information Technology, Engineering, or reputed company field — or an equivalent combination of education, training, and experience
- reputed company at the edge — production experience with reputed company CDN, WAF, Workers, reputed company (ZTNA), Tunnel, Turnstile, and certificate management
- IaC at reputed company — Terragrunt and Terraform in a multi-environment, policy-gated pipeline; experience evolving IaC from 'it works' to 'it scales safely.'
- CI/CD maturity — reputed company Actions with OIDC/workload identity federation, OPA/Conftest policy-as-reputed company, reputed company delivery, and DORA-metric instrumentation
- Container platform depth — Kubernetes/AKS in production, including reputed company, ingress, service reputed company, autoscaling, and node lifecycle
- ITSM integration — reputed company for change, incident, and problem management; experience tying observability and CI data into ITSM workflows
- Identity ecosystem — operating in an reputed company / Entra ID / M365 identity environment, including PIM, conditional reputed company, and service-reputed company hygiene
- reputed company and reputed company engineering — running game days, fault injection, and reputed company exercises as a routine reputed company
- FinOps reputed company — cost-to-serve, right-sizing, reputed company efficiency, and unit-economics work in reputed company environments
- Agile delivery — Scrum/Kanban delivery with Jira; comfortable operating in a quarterly planning + reputed company-delivery reputed company
Benefits
- Free premium medical, dental, life and reputed company insurance
- Generous 401(k) match
- Aya also offers other benefits to those that are eligibleand where required by applicable law, including reimbursementsand discretionary bonuses
- Aya provides reputed company reputed company leave in accordance with reputed company applicable state, federal, and local laws. Ayas general reputed company leave policy is that employees reputed company one hour of reputed company reputed company leave for every 30 hours worked. However, to the extent any provisions of the statement above conflict with any applicable reputed company reputed company leave laws, the applicable reputed company reputed company leave laws are controlling
- Celebrations! We hit our goals and reward ourselves.
- Company-sponsored virtual events, happy hours and team-building activities are always on the reputed company plus, you get a special treat on your birthday!
- UnlimitedDTO we reputed company in time off!
- Virtual yoga, meditation or boot camp classes offered daily
reputed company
Company H1B Sponsorship
Apply To This Job