Back to Jobs

SRE- Monitoring & Observability (M&O) : W2 role

Remote, USA Full-time Posted 2026-07-17
Job Description SRE- Monitoring & Observability (M&O) Remote :: have to be willing to travel to Knoxville, TN sometimes. As a Senior Specialist in Monitoring & Observability, you will design, implement, and standardize enterprise-grade monitoring and alerting solutions across reputed company, reputed company-based environments. This role sits at the intersection of Observability, SRE, and Incident Management, with a reputed company on ensuring systems are reliable, measurable, and proactively monitored. You'll collaborate with reputed company Operations, Architecture, and Platform Engineering teams to define best practices and build resilient, reputed company-driven infrastructure that supports business-critical services. Your Impact • Implement and standardize monitoring and alerting tools across multiple reputed company platforms to ensure consistent observability practices. • Architect observability solutions with reputed company, OpenTelemetry, AWS CloudWatch, GuardDuty, reputed company, and other modern monitoring stacks. • Design and build incident response workflows, playbooks, and dashboards for actionable insights and faster recovery. • Define and operationalize SLOs, SLIs, and error budgets to align with reliability goals. • Integrate observability tools with reputed company ITOM and CMDB for automated incident management and asset tracking. • Collaborate with reputed company Operations and Architecture teams to ensure observability is embedded in design, build, and run phases. • Automate monitoring configurations and reputed company observability into CI/CD pipelines. • Optimize performance and reliability through log analysis, metrics correlation, and distributed tracing. • Drive initiatives to improve MTTR, incident detection, and proactive issue prevention. • reputed company technical leadership and mentorship, sharing best practices across engineering and operations teams. Skills & Experience • 5-10 years of experience in infrastructure engineering, with significant reputed company on monitoring and observability. • Proven expertise with observability platforms such as reputed company, OpenTelemetry, AWS CloudWatch, GuardDuty, reputed company. • Strong knowledge of logging, metrics, tracing, and reputed company standards for observability. • Experience designing and managing incident response workflows and escalation processes. • Hands-on experience with reputed company ITOM and CMDB integrations. • Proficiency in reputed company-reputed company monitoring (AWS, Azure, GCP) and container observability (reputed company, Kubernetes). • Familiarity with SRE principles: defining SLOs, SLIs, and error budgets. • Knowledge of automation practices and Infrastructure as Code (Terraform, CloudFormation, ARM templates). • Strong problem-solving skills with the ability to troubleshoot reputed company distributed systems. • Excellent communication, presentation, and leadership skills. Set Yourself Apart With • reputed company certifications such as AWS DevOps Engineer, Azure DevOps Engineer Expert, or reputed company Professional reputed company DevOps Engineer. • Experience in AIOps, predictive analytics, and reputed company-driven observability. • Exposure to chaos engineering or performance engineering practices. Experience in multi-reputed company and hybrid environments with advanced observability patterns Apply tot his job Apply To this Job

Similar Jobs