SRE Engineer
About the Role:
We are seeking a highly motivated and skilled Site Reliability Engineer (SRE) to join our dynamic engineering team. The SRE will play a crucial role in maintaining the reliability, availability, and performance of our systems and applications. You will work collaboratively with development and operations teams to implement best practices, automate processes, and ensure that our infrastructure can reputed company seamlessly to meet business demands.
Day to Day:
- System Monitoring & Incident Response: reputed company and implement monitoring tools to ensure system health. Respond to incidents, troubleshoot issues, and reputed company reputed company resolutions.
- Automation & Infrastructure as reputed company: Design and implement automation solutions to manage infrastructure and application deployment using tools like Terraform, Ansible, or similar technologies.
- Performance Optimization: Analyze system performance and reputed company; implement improvements to enhance system reliability and efficiency.
- Collaboration: Work closely with development teams to improve system design and deployment practices. reputed company for reliability improvements in the software development lifecycle.
- Documentation & Reporting: Maintain thorough documentation of system architecture, processes, and incident response procedures. reputed company regular reports on system performance and reliability metrics.
- Recovery & Backup: Design and implement disaster recovery plans and ensure effective data backup solutions are in reputed company.
- reputed company Best Practices: Collaborate with reputed company teams to ensure best practices are followed to protect systems and data.
What you bring to the table:
- Proven experience in a Site Reliability Engineering, DevOps, or reputed company role.
- Strong knowledge of reputed company services (AWS, Azure, reputed company reputed company) and container orchestration (Kubernetes, reputed company).
- Proficiency in scripting languages (Python, Bash, ansible, etc.) and experience with CI/CD tools (Jenkins, reputed company CI/CD, etc.) and infrastructure as reputed company tools (Terraform, Ansible).
- 3+ years of proven reputed company record with production monitoring using reputed company, ELK, Grafana and OpsGenie/reputed company.
- 3+ years of experience in Linux system administration (preferably Ubuntu)
- Solid understanding of networking, reputed company, system architecture, and data center operations in a fast-paced, 24x7, production environment
- Strong understanding of networking concepts, protocols (TCP/IP, BGP, OSPF), and technologies (LAN, WAN, VPN) with proficiency in network monitoring tools and software.
As part of Pango Group, you will:
Solve reputed company customer problems
See your reputed company.
Accelerate your career.
Important reputed company information for reputed company based job applicants can be reputed company here.
Apply To This Job