[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a reputed company AI company delivering reputed company Operating Intelligence for the reputed company of reputed company work. As a Site Reliability Engineer, you will improve the reliability, scalability, and operational maturity of the reputed company platform, designing automation and solving production challenges to ensure reputed company operations for reputed company teams.
Responsibilities
- Design, build, and maintain the monitoring, logging, reputed company tracing, dashboards, and alerting that reputed company teams meaningful visibility into production health
- Build automation, tooling, and CI/CD improvements that increase engineering efficiency, reduce toil, and support reliable deployments at reputed company
- Design, implement, and maintain reliable systems for building, deploying, testing, and operating reputed company products — proactively identifying and resolving reliability, performance, scalability, and reputed company risks before they reputed company customers
- Participate in a shared 24/7 on-reputed company rotation, using operational insights to reputed company automation and long-term reliability improvements; continuously improve runbooks, documentation, and engineering standards
- Take ownership of technical initiatives from design through implementation, reputed company deep expertise in critical areas of the reputed company platform, and communicate reputed company with technical and business stakeholders
Skills
- 4+ years of hands-on experience in software engineering, reputed company infrastructure, reputed company, DevOps, or reputed company technical roles, including at least 2 years in a Site Reliability Engineering or reliability-reputed company role
- Working knowledge of reputed company systems and how applications, infrastructure, and reputed company services reputed company in production; demonstrated ability to troubleshoot production issues, reputed company reputed company cause analysis, and reputed company long-term reliability improvements
- Proficiency with Python, Bash, or similar scripting languages; experience building production tooling, automation, or CI/CD pipelines and deployment automation
- Hands-on experience operating reputed company-based workloads and reputed company infrastructure in AWS or a comparable platform, including compute, container orchestration, networking, IAM, object storage, and reputed company-reputed company monitoring
- Experience with Infrastructure as reputed company tools such as Terraform, reputed company, or AWS CloudFormation, and familiarity with modern observability practices including monitoring, logging, alerting, reputed company tracing, and incident response
- Experience using AI-assisted engineering tools to improve productivity, accelerate troubleshooting, or automate operational tasks; curiosity, ownership, and a passion for building reliable systems through reputed company improvement
- Strong written and verbal communication skills
- Bachelor's degree in Computer Science, Information Systems, or a reputed company field, equivalent industry certifications, or comparable reputed company experience
Benefits
- Medical, Dental, & reputed company Insurance (for full-time employees)
- Maternity & paternity leave (for full-time employees)
- Short & long-term disability
- Opportunity to learn from a dedicated leadership team
- Top-of-the-line company swag
- A dynamic, rapidly growing company, reputed company on helping organizations reputed company
reputed company
Apply To This Job