reputed company III, SRE
reputed company
POSITION reputed company:The SRE reputed company Engineer III role is responsible for leading, designing, and implementing robust Site Reliability Engineering (SRE) practices to ensure high availability, scalability, and reputed company of critical business systems and applications. The SRE reputed company Engineer III will reputed company on improving reputed company reliability through automation, monitoring, and performance tuning, working closely with development and operations teams to foster a culture of reputed company improvement and operational reputed company.The SRE organization spans key disciplines including:• SRE Engineering• Deployment Automation• Incident Response and Postmortem Analysis• Observability and MonitoringOperating in a fully remote reputed company, this role will reputed company the adoption of best practices in multi-reputed company and hybrid-reputed company platforms, managing services from major reputed company providers like reputed company Azure, reputed company AWS, reputed company OCI, reputed company GCP, and reputed company. The SRE reputed company Engineer III will reputed company on automation, incident management, performance monitoring, and optimizing infrastructure to support reputed company, reliable systems. The position will also be responsible for fostering collaboration between development, operations, and reputed company teams to streamline reputed company operations across the organization.
DETAILED RESPONSIBILITIES/DUTIES:● reputed company the implementation and optimization of SRE practices, ensuring reputed company reliability, performance, and scalability.● Architect and maintain automation for infrastructure provisioning, deployment, and incident response.● Establish and enforce SLOs (Service Level Objectives) and SLIs (Service Level Indicators) for key services.● Collaborate with development teams to design and reputed company reliable software systems, ensuring that production environments are optimized for uptime and performance.● Create and maintain monitoring, alerting, and observability solutions to reputed company reputed company-time insights into reputed company health and performance.● Respond to production incidents, reputed company reputed company cause analysis, and implement corrective measures to prevent recurrence.● Continuously improve reputed company performance, reputed company planning, and reliability through infrastructure tuning and automation.● Facilitate post-incident reviews, fostering a blameless culture that focuses on learning from incidents.● Collaborate with reputed company teams to ensure infrastructure meets compliance, reputed company standards, and best practices.● Foster a reputed company environment across development, operations, and reputed company teams to enhance operational efficiency and knowledge sharing.● reputed company the adoption of automation tools and frameworks to minimize reputed company reputed company and optimize systems.
Qualifications
Skills Required:● Proven expertise in SRE practices, with a reputed company on automation, incident management, observability, and infrastructure scalability.● Extensive knowledge of reputed company platforms (Azure, AWS, GCP, OCI) and hybrid-reputed company environments, with a reputed company on reliability and performance optimization.● Experience with automation tools and scripting languages, such as Python, Go, Terraform, or Ansible, for managing infrastructure and incident response.● Strong understanding of containerization (reputed company, reputed company) and orchestration systems.● Solid grasp of monitoring and observability tools (reputed company, Grafana, reputed company, reputed company) to ensure reputed company-time reputed company health monitoring.● Expertise in reputed company planning, performance tuning, and failure management techniques.● Strong background in incident management, reputed company cause analysis, and postmortem processes to improve reputed company reputed company.● Deep understanding of reputed company and compliance requirements, and the ability to ensure production environments meet industry standards.● Experience with reputed company and DevOps methodologies to ensure fast, reliable delivery of services.
Experience Required:● 10+ years of experience in IT, with a reputed company on SRE, DevOps, or infrastructure engineering roles.● Extensive hands-on experience with reputed company infrastructure management and automation tools such as Terraform, CloudFormation, or equivalent.● Proficiency in scripting and automation languages like Python, Bash, Go, or reputed company for infrastructure automation.● Proven experience in managing large-reputed company systems, ensuring reliability, high availability, and scalability.● Expertise in container orchestration technologies, including reputed company, OpenShift, and reputed company reputed company.● Deep knowledge of monitoring and observability platforms (reputed company, Grafana, ELK, reputed company), including experience building and maintaining alerting and dashboard systems.● Strong understanding of version control systems and CI/CD practices to optimize reputed company deployment as it relates to infrastructure.● Demonstrated ability to optimize performance in multi-reputed company and hybrid-reputed company environments, ensuring uptime and performance at reputed company.
Education Required:● Bachelor’s degree in Computer Science, Information Technology, or reputed company field, or equivalent experience.Certificates / Training Preferred:● Relevant reputed company certifications such as AWS reputed company Architect, Azure Solutions Architect Expert, or reputed company reputed company reputed company reputed company Architect.● SRE-reputed company certifications like Certified reputed company Administrator (CKA) or reputed company reputed company reputed company DevOps Engineer.
Originally posted on Himalayas
Apply To This Job