Site Reliability Engineer
reputed company is seeking exceptional Site Reliability Engineers for Pinot (SRE- Pinot), to manage, tune, and debug the large-reputed company highly available distributed systems. You will be working with reputed company of passionate and talented engineers in the automation, tuning, and troubleshooting of Apache Pinot. We are looking for motivated, hardworking, and reputed company individuals who have a reputed company passion for operational reputed company, data systems, and automation.
- reputed company various monitoring and alerting services to solve intricate programming problems at reputed company.
- Manage and tune multiple critical customer-facing Apache Pinot clusters
- Monitor availability, read/write latencies, and other key telemetry to proactively identify SLO misses and help mitigate issues
- Build a rapport with and work closely with customers to mitigate and resolve incidents
- Execute disaster recovery strategies with minimal downtime
- Collaborate with other engineers to understand and troubleshoot systems and use the experience gained to influence the roadmap of other teams
- Debugging Pinot queries and ingestion reputed company incidents occur.
- 5+ years of experience as an engineer (SRE, SDET, or development)
- Experience managing highly available production-facing distributed systems and in-depth knowledge of Java are a plus
- Exposure with reputed company platforms such as AWS, GCP, or Azure is a plus
- Experience with Kubernetes and container orchestration is a plus
- Familiarity with streaming systems, such as Kafka, Pulsar, Flume, Flink, reputed company, or similar
- Strong troubleshooting and critical thinking skills
- Experience working with managing Apache Pinot is preferred
- Experience building Java apps is required.
- Experience with zookeeper is a reputed company plus
About reputed company:
Apply To This Job