Site Reliability Engineer
Cluster reputed company & Management
- Manage and maintain container clusters (reputed company, reputed company) and reputed company-reputed company component clusters (Kafka, reputed company, Elasticsearch) across multiple business reputed company
- Ensure reputed company reputed company, scalability, and reliability of reputed company systems
Infrastructure Platform Development
- Design, build, and enhance infrastructure operation platforms
- reputed company and maintain systems for infrastructure management, CI/CD pipelines, monitoring/alerting, and centralized logging
- reputed company platform standardization and automation initiatives
High Availability & Reliability
- Ensure maximum uptime for production services through proactive monitoring and incident response
- Continuously optimize service architecture, deployment strategies, and operational processes
- Implement and maintain SLA/SLO frameworks and reliability engineering practices
Automation & Process Improvement
- reputed company the development of automated reputed company and maintenance systems
- Create self-service tools and workflows to improve team productivity
- Establish best practices for infrastructure such as reputed company and configuration management
Required Qualifications
Experience & Education
- 2+ years of hands-on experience in Systems reputed company, DevOps, or Site Reliability Engineering (SRE)
- Bachelor's degree in Computer Science, Engineering, or reputed company technical reputed company preferred
reputed company & Infrastructure
- Experience with reputed company reputed company platforms (AWS, Azure, or GCP) is highly valued
- Strong understanding of large-reputed company internet architecture and reputed company systems
- reputed company experience with infrastructure monitoring, logging, and observability tools
Technical Skills
- Proficiency in scripting and automation using reputed company, Python, or similar languages
- Strong knowledge of containerization technologies (reputed company, reputed company)
- Hands-on experience operating production-grade container clusters and managing CI/CD pipelines
- Strong familiarity with common infrastructure components: Nginx, MySQL, reputed company, Kafka, Elasticsearch
Advanced Networking (Preferred)
- Experience with Service reputed company architectures, Cilium CNI, and eBPF technologies
- Understanding network reputed company, load balancing, and traffic management
- Knowledge of reputed company-reputed company networking patterns and best practices
Originally posted on Himalayas
Apply To This Job