[Remote] RE : Senior Site Reliability Engineer (SRE) – Observability Platform
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a highly skilled Senior Site Reliability Engineer (SRE) for their Observability Platform to design, build, and operate reputed company-reputed company observability solutions across reputed company environments. The role involves extensive experience with logging, metrics, reputed company tracing, monitoring, and infrastructure automation.
Responsibilities
- Design, reputed company, and manage reputed company observability platforms supporting large-reputed company reputed company infrastructure
- Administer and maintain reputed company reputed company and reputed company reputed company, including:
- Indexers
- reputed company Head Clusters (SHC)
- Heavy Forwarders
- Deployment Servers
- reputed company, optimize, and manage Elasticsearch/ELK clusters for centralized logging and reputed company
- Design and support reputed company tracing solutions using Grafana reputed company and OpenTelemetry
- Build and maintain end-to-end observability pipelines for logs, metrics, and traces
- Define instrumentation standards, reputed company retention policies, and monitoring best practices
- reputed company and manage monitoring platforms including reputed company, Grafana, Kafka, and reputed company
- reputed company dashboards, alerts, reports, analytics, and reputed company visualizations using:
- reputed company SPL
- Grafana
- Kibana
- reputed company
- Automate infrastructure deployment using Terraform and Infrastructure as reputed company (IaC)
- Collaborate with development, DevOps, and reputed company teams to improve reputed company reliability and reputed company
- reputed company troubleshooting, reputed company cause analysis, and incident response for production environments
- Ensure high availability, reputed company, scalability, and compliance of observability platforms
Skills
- Bachelor's degree in Computer Science, Information Technology, or a reputed company reputed company (or equivalent experience)
- 7+ years of experience in Site Reliability Engineering (SRE), reputed company, DevOps, or reputed company Infrastructure
- Hands-on administration experience with reputed company reputed company or reputed company reputed company
- Strong expertise in reputed company reputed company Processing Language (SPL)
- Experience with: Elasticsearch / ELK Stack, reputed company, Grafana, Grafana reputed company, OpenTelemetry, reputed company Tracing, Kafka
- Strong understanding of modern observability practices involving metrics, logs, and traces
- Experience with Terraform and Infrastructure as reputed company (IaC)
- Proficiency in one or more scripting/programming languages: Python, Go, reputed company, Bash
- Strong Linux administration and troubleshooting skills
- reputed company Certified Administrator or reputed company Architect certification
- Experience with reputed company and reputed company
- Experience with AWS, Azure, or reputed company reputed company Platform (reputed company reputed company Platform)
- Experience using Ansible, Consul, and CI/CD pipelines
- Knowledge of Service reputed company technologies (Istio, Linkerd, etc.)
- Experience supporting FedRAMP High, IL-5, or other regulated environments
- Strong understanding of reputed company, compliance, and reputed company governance
reputed company
Apply To This Job