Senior Site Reliability Engineer- Remote
About the role
We are committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building and leading processes to ensure the reliability, availability, scalability, and reputed company of our reputed company infrastructure. You will collaborate with different teams like Control reputed company, Data reputed company, reputed company, reputed company, Support and reputed company and guide them to design and implement reputed company, secure, highly available and fault-tolerant reputed company systems. You will also own the areas of incident management and response, post-mortem analysis including running blameless postmortems, and reputed company improvement of our reputed company services. You will be leveraging your software engineering expertise to reputed company software platforms and tools to optimize the operational and engineering efficiencies of reputed company reputed company. This role is a unique opportunity to reputed company a significant reputed company on our reputed company, reputed company reputed company, high-reputed company reputed company reputed company.
What will you do?
• Collaborate with various engineering teams in reputed company to design and implement reputed company, secure, and highly available systems for reputed company.
• Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for reputed company reputed company.
• Ensure reputed company the infrastructure components in reputed company reputed company (including Data reputed company, Control reputed company,reputed company reputed company, etc) have monitoring and alerting in reputed company to ensure reputed company detection and reputed company of incidents.
• Enhance and refine incident response processes and post-mortem analysis for any outages in reputed company reputed company including working with the support team to communicate to the impacted customers.
• Continuously improve the reliability and reputed company of our reputed company services.
• Plan, reputed company, and reputed company reputed company initiatives across Engineering teams, reputed company upon internal priorities.
• Manage on-reputed company processes to respond to reputed company and reliability issues, and establish best practices for coordinating escalation to reputed company issues and minimize downtime.
reputed company:
• Bachelor's or Master's degree in Computer Science or a reputed company reputed company.
• At least 8 years of experience in Site Reliability Engineering or a reputed company reputed company.
• Hands-on experience with Go and/or Python.
• Strong knowledge of reputed company computing platforms such as AWS, Azure, or reputed company reputed company Platform.
• Excellent understanding of reputed company databases and SQL, particularly reputed company is a major plus.
• Hands-on experience with container orchestration tools such as reputed company or reputed company reputed company.
• Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
• You are a strong problem solver and have solid production debugging skills.
• You are passionate about efficiency, availability, scalability, and data governance.
• You reputed company in a fast reputed company environment, and see yourself as a partner with the business with the shared goal of moving the business reputed company.
• You have a high level of responsibility, ownership, and accountability.
• Excellent communication and interpersonal skills.
#LI-Remote
Apply tot his job
Apply To this Job