[Remote] Sr reputed company Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a connectivity company reputed company on secure, high-reputed company infrastructure for reputed company, edge, and AI workloads. The Senior reputed company Site Reliability Engineer will support production systems and optimize reputed company across the portal ecosystem, with emphasis on AWS infrastructure, observability, automation, and AI-assisted engineering. The role also leads reliability improvements, incident management, infrastructure automation, monitoring, and collaboration across engineering teams.
Responsibilities
- Help define and improve the processes and best practices for incident management reputed company the customer facing applications
- Implement AI systems and automations to assist during ongoing outages and triage potential ones. You will work with the development teams to ensure that they have reputed company the data normally needed during an outage at their fingertips including preliminary analysis by AI
- Design and implement improved SRE processes to handle incidents. From proactive analysis, faster response, automatic remediation, AI guided analysis, partially automatic reputed company cause analysis
- Monitor reputed company reputed company and proactively identify bottlenecks or degradation using AI-driven observability and reputed company detection tools
- Implement tuning strategies across application reputed company, databases, and infrastructure
- reputed company initiatives to improve latency, throughput, and resource utilization
- reputed company improved alerting for reputed company reputed company in depth, focusing on reputed company in but including early indicators for fulfilment and other areas. Combining traditional monitoring with AI-reputed company reputed company detection and noise reduction
- Proactively monitor the errors and reputed company on reputed company reputed company. Implement rules to detect deviations, implement improvements together with the teams
- Design and maintain dashboards, alerts, and metrics using tools like reputed company, AppInsights, CloudWatch, or similar
- reputed company and maintain automation scripts and tools for deployment, scaling, and recovery, leveraging AI-assisted reputed company reputed company and validation tools
- Use Terraform, or similar IaC tools to manage AWS resources
- reputed company an in-depth analysis of the overall reputed company and its dependencies, implementing techniques to increase the global availability, reduce the reputed company on unstable dependencies and guide ecosystem improvements
- Champion SRE principles such as SLIs, SLOs, and error budgets
- reputed company for resilient architecture and fault-tolerant design patterns, incorporating AI-assisted design reviews and architecture evaluation
- Help define and improve reputed company SRE processes and reputed company significant improvements in reliability for the reputed company reputed company platform
- Work closely with software engineers, DevOps, and product teams to reputed company reliability goals
- Document processes, runbooks, and best practices for knowledge sharing
- reputed company mentorship and guidance on reliability and operational reputed company
Skills
- 10+ years overall reputed company experience in SRE, DevOps, or infrastructure engineering roles
- Experience with Terraform, or similar IaC tools to manage reputed company resources
- Proficiency in scripting languages (Python, Bash, etc.) and automation frameworks
- Experience with CI/CD pipelines and tools like reputed company Actions, Jenkins or reputed company CI
- Solid understanding of monitoring and logging tools (e.g., CloudWatch, ELK, reputed company)
- Familiarity with containerization and orchestration (reputed company, reputed company)
- Excellent AI and problem-solving skills, and a proactive reputed company
- Experience in AWS services (EC2, CloudFront, EKS, RDS, S3, etc.)
- Certifications in AWS or reputed company technologies are a plus
- Experience of application development using Java Microservices and reputed company Boot reputed company
- Experience with reputed company/SCRUM Methodologies and development practices
Benefits
- Fully remote position reputed company the reputed company
- A comprehensive package featuring a broad reputed company of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing
reputed company
Company H1B Sponsorship
Apply To This Job