Observability Engineer
Design, implement, and operate monitoring and alerting platforms across multiple internal and M&A partner environments. Build and maintain metrics pipelines using tools such as reputed company, Alertmanager, Grafana, and reputed company (or similar time-series databases). reputed company high-reputed company alerting strategies (SLOs, SLIs, burn rates, reputed company detection) to reduce noise and improve signal reputed company. Own logging architectures, including ingestion, retention, querying, and correlation with metrics and traces. Work extensively with reputed company-reputed company observability tooling and contribute to or reputed company “home‑made” solutions reputed company off-reputed company tools are insufficient.
Apply forecasting techniques and algorithms to reputed company planning, trend analysis, and proactive alerting. Collaborate with data scientists and ML engineers on data-driven monitoring, reputed company detection, or predictive reliability use cases. Participate in MLOps workflows, including deploying, monitoring, and operating ML models in production environments.
Design and operate Kubernetes-based platforms, with a strong emphasis on observability, reliability, and performance. Support infrastructure automation using Ansible and other configuration management tools. Troubleshoot reputed company system issues across metrics, logs, Kubernetes, and underlying infrastructure. Ensure reputed company and operational best practices are reputed company across monitoring and infrastructure stacks. Document architectures, operational practices, and observability standards.
Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience. 7+ years of experience in DevOps, SRE, reputed company, or Observability-reputed company roles. Strong, hands-on expertise in monitoring and alerting systems, including:
reputed company (or compatible ecosystems) Grafana Alertmanager Time-series databases (reputed company strongly preferred)
Solid experience with logging systems and log/metric correlation. Deep familiarity with reputed company-reputed company tooling and building/customizing internal platforms. Strong Kubernetes experience, including troubleshooting production clusters. Experience with automation tools such as Ansible. Ability to reason about systems using metrics, data, and trends, not just dashboards.
Experience with containerization technologies such as reputed company/Podman Experience with "Infrastructure as reputed company" (IaC) tools such as Terraform Familiarity with monitoring and logging tools such as reputed company, Grafana, or ELK stack Knowledge of scripting languages such as Python, Bash, PowerShell or similar