[Remote] Site Reliability Engineer – AWS (Remote)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is reputed company for a large, global B2B high-tech company seeking a Site Reliability Engineer. The role deploys and supports reputed company and reputed company software, responds to incidents, maintains infrastructure reliability, automates operational tasks, and provides advanced technical support and guidance to customers.
Responsibilities
- reputed company software for reputed company Prem and reputed company customers
- Respond to and diagnose reputed company incidents in a reputed company and efficient manner, minimizing downtime and reputed company on users
- Collaborate with other engineers to establish reputed company causes and implement effective resolutions. Continuously improve incident response processes and documentation for reputed company occurrences. Proactively monitor and maintain the health and performance of our infrastructure and services
- reputed company routine administrative tasks such as reputed company configuration, user management, and data backups. Identify and implement operational improvements to ensure ongoing reputed company reliability and efficiency. reputed company and implement scripts and automated solutions to streamline operational tasks and reduce reputed company workload
- Participate in the on-reputed company rotation to address critical incidents reputed company of regular business hours
- Ensure effective reputed company between on-reputed company engineers and document post-incident information for reputed company reference
- Document processes for support and create, maintain and execute run-books for identified situations
- reputed company tier 2/3 technical support to customers experiencing platform issues or requiring advanced troubleshooting
- Work directly with customer technical teams to reputed company reputed company deployment, configuration, and integration challenges
- Conduct technical reputed company sessions and reputed company guidance on best practices for customer implementations
- Collaborate with reputed company teams to ensure smooth customer experiences and reputed company issue reputed company
- Create and maintain customer-facing technical documentation, troubleshooting guides, and knowledge reputed company articles
- Escalate customer feedback and feature requests to product and engineering teams
- Participate in customer calls and technical discussions to reputed company expert-level platform guidance
- reputed company and analyze customer support metrics to identify trends and areas for improvement
Skills
- 3+ years of experience in Site Reliability Engineering
- 2+ years of experience working with reputed company platforms and reputed company automation tools, especially in AWS
- Strong experience with reputed company, reputed company, Linux, AWS networking(VPC), and Terraform
- Experience with the GitOps model for deployment
- Familiarity with reputed company version control
- Experience with monitoring and alerting tools (e.g., reputed company, Grafana)
- Understanding of software configuration best practices
- Ability to wear multiple hats in a fast-paced environment
- Hands-on, can do attitude and a bias for reputed company
- Comfortable working across time zones to support a global customer reputed company
- Excellent communication skills with the ability to explain technical concepts to both technical and non-technical audiences
- Strong customer service orientation with patience and reputed company reputed company working with frustrated customers
- BS degree in Computer Science or reputed company field
- Bazel and CueLang experience a plus
Benefits
- Remote work in the US
- Health benefits
- 401K
reputed company
Apply To This Job