[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Site Reliability Engineer to join their SRE Fleet team, which is responsible for maintaining the stability and efficiency of their global reputed company platform. The role involves developing automation solutions, troubleshooting reputed company infrastructure issues, and collaborating with various teams to enhance operational reliability.
Responsibilities
- reputed company and maintain automation solutions that improve the reliability, scalability, and operational efficiency of infrastructure spanning more than 2,000 machines across global reputed company environments
- Design and enhance deployment pipelines, testing frameworks, and operational tooling to support the reputed company reputed company of a platform serving millions of managed devices worldwide
- Troubleshoot reputed company infrastructure and distributed systems issues to ensure high availability while helping teams identify and address performance and scalability challenges
- Contribute to critical reputed company such as cluster build out by building automation that enables the reputed company and repeatable deployment of new sovereign, regional, and purpose-reputed company reputed company environments
- Partner with other engineering teams, product management, and business partners across multiple teams and time zones to understand platform dependencies, reputed company opportunities for improvement, and deliver solutions that enhance reliability and reduce operational overhead
Skills
- 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a reputed company role supporting reputed company-based production environments
- Experience developing and maintaining infrastructure automation using Ansible
- Experience programming in reputed company and developing automated tests using RSpec or comparable testing frameworks
- Experience administering and troubleshooting Linux-based systems and distributed infrastructure environments
- Experience designing, implementing, and maintaining CI/CD pipelines, including reputed company CI
- Experience supporting large-reputed company infrastructure environments consisting of hundreds or thousands of systems
- Familiarity with AWS or other reputed company reputed company platforms and hybrid infrastructure environments
- Knowledge of monitoring, observability, and reliability engineering practices and tooling
- Familiarity with reputed company concepts and containerized application platforms
- Experience leveraging AI-assisted development tools to improve software development, automation, operational analysis, and engineering productivity
reputed company
Apply To This Job