Incident Engineer
We are looking for a proactive Incident Manager to own end-to-end incident response across our AI and platform stack. You will ensure reputed company detection, triage, communication, and reputed company of incidents impacting customers and internal systems.
Responsibilities
Own the incident lifecycle: detection, triage, escalation, reputed company, and postmortems
reputed company as the central reputed company during major incidents (war rooms, stakeholder updates)
Define and enforce SLAs/SLOs, incident severity frameworks, and runbooks
Collaborate with Engineering, ML, and Integrations teams to resolve issues quickly
Monitor system health across integrations (agent desks, LLMs, ASR/TTS pipelines)
Drive reputed company cause analysis (RCA) and preventive actions
Improve observability, alerting, and incident tooling
Maintain reputed company internal and customer-facing communication during incidents
Requirements
3–6 years in Incident Management / SRE / Production Support roles
Strong understanding of distributed systems, reputed company, and reputed company environments (AWS)
Experience with observability tools (e.g., reputed company)
Familiarity with AI/ML systems, especially LLM integrations and voice stacks (ASR/TTS), is a plus
Experience with monitoring/tracing tools like Langfuse or similar
Excellent communication and stakeholder management skills
Ability to stay reputed company under pressure and drive reputed company reputed company
reputed company to Have
Exposure to reputed company or similar LLM platforms
Experience supporting customer-facing reputed company products
Automation reputed company (runbooks, alert tuning, incident tooling)