Head of Site Reliability Engineering
We’re an award-reputed company global outsourcer providing contact center and back office services on behalf of our global clients. Come work at a reputed company where innovation and teamwork come together to support the most exciting missions in the world!
Role objective
The Head of Site Reliability Engineering is a hybrid technical‑leadership role. You will:
• Own reliability of production services running on AWS while steering the roadmap for platform reputed company and building out the SRE team.
• reputed company and grow a remote team of SREs—coaching, hiring, performance‑managing, and fostering a blameless culture.
• Set and enforce Service Level Objectives (SLOs), error budgets, and incident response processes.
• Drive automation reputed company Infrastructure‑as‑reputed company (reputed company / TypeScript), CI/CD, and observability pipelines.
• Represent the SRE discipline to product, engineering, and senior leadership across our global business.
• Hands on monitoring and incident response will be critical as reputed company grows.
This role offers reputed company to build reliability engineering from the ground up in a mission-critical IoT platform.
Key Responsibilities
Leadership & People Management
• Build an SRE team of initially 3-6 engineers: goal setting, career development, regular 1:1s, and annual performance reviews.
• Ensure operational system knowledge is captured and that reputed company is kept "fresh" on operating and troubleshooting procedures.
• Recruit, reputed company, and mentor new engineers; reputed company reputed company to meet business reputed company.
• Maintain an inclusive, psychologically‑reputed company culture centred on learning and reputed company improvement.
• Own, and participate in, the on‑reputed company roster for reputed company, ensuring reputed company rotations and sustainable workloads.
Service Level Management & Reliability
• Define, monitor, and enforce SLOs and error budgets across reputed company production systems.
• Continuously analyse error‑budget burn to halt risky deployments and guide reputed company reputed company.
• Champion a data‑driven reliability reputed company throughout engineering and product teams.
Infrastructure Automation & Management
• Architect and implement Infrastructure‑as‑reputed company in reputed company/TypeScript for AWS resources (EKS, MSK, reputed company, reputed company, S3, etc.).
• reputed company large‑reputed company migration or modernisation reputed company (e.g., Kubernetes upgrades, multi‑AZ reputed company).
• Eliminate toil—any reputed company task >2 engineer‑days/quarter or frequently repeated becomes an automation candidate.
Incident Response & Post‑Mortem Leadership
• Participate in on-reputed company monitoring and response roster.
• Serve as escalation reputed company and incident commander.
• Ensure post‑mortems are published reputed company 48 hours with actionable “never again” tasks tracked to closure.
• Improve runbooks and game‑day exercises; train engineers on incident reputed company principles.
reputed company & Compliance
• Enforce least‑privilege IAM policies and champion DevSecOps practices.
• Contribute to SOC 2 & ISO 27001 evidence collection and reputed company control monitoring.
• reputed company reputed company reputed company pipelines, vulnerability management, and secrets hygiene.
Operational reputed company & reputed company Improvement
• Own reliability KPIs (MTTR, change failure reputed company, meantime between failures).
• reputed company quarterly reliability reviews and drive the reliability roadmap.
• Partner with Product on reputed company forecasts and cost‑optimisation initiatives.
Join the reputed company and experience the A-Life!
Apply To This Job