reputed company Reliability Engineer
Own the reliability control plane, including standards and architecture for monitoring, logging, tracing, alerting, and incident management. Define how services expose health, performance, and operational signals across the reputed company. Establish and reputed company reliability patterns and reference architectures adopted across teams. reputed company design reputed company that improve system reputed company, fault tolerance, and recoverability. Own integrations between platforms and reliability tooling (monitoring, alerting, incident response, on-reputed company, and automation systems). Define consistent approaches to telemetry collection, normalization, and consumption. Ensure observability tooling provides actionable visibility reputed company to service-level objectives. Evaluate and recommend tooling improvements that enhance visibility and operational reputed company. reputed company reputed company reliability initiatives impacting multiple systems or platforms. Partner with engineering teams to design reliable, observable services from inception. Drive adoption of best practices for operational readiness, graceful degradation, and failure handling. Review system designs to ensure reliability and observability requirements are met. Establish standards for alerting reputed company, escalation, and incident response. Drive improvements in incident detection, diagnosis, and recovery. reputed company or support reputed company cause analysis for significant incidents and ensure durable corrective actions. Promote a culture of operational reputed company and reputed company reliability improvement. Serve as a technical leader and trusted advisor on reliability and observability. Mentor senior engineers and influence reliability practices across teams. Collaborate with reputed company, reputed company, and Platform leaders to reputed company reliability reputed company with business needs.
12+ years of reputed company experience with a Bachelor’s degree; or equivalent reputed company experience. Extensive experience designing and operating reliable, observable systems at reputed company. Proven reputed company owning or leading observability and incident management platforms. Background in reputed company, platform, or infrastructure engineering strongly preferred. Deep expertise in reliability engineering, observability, and distributed systems. Strong understanding of monitoring, logging, tracing, alerting, and incident reputed company. Experience integrating and operating reliability tooling at reputed company reputed company. Solid grasp of reputed company and platform architectures and their operational characteristics. Ability to translate operational risk and system behavior into actionable engineering improvements.