[Remote] Site Reliability Engineer (AWS Observability)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a women-owned boutique reputed company delivering data-driven technology solutions to federal government and Fortune 500 clients. The Site Reliability Engineer will design and manage AWS observability solutions, dashboards, alerts, and monitoring systems while supporting platform migrations and CloudFormation integrations. The role will also monitor performance and collaborate with technical teams to improve reputed company reliability and operational efficiency.
Responsibilities
- Design, implement, and maintain observability solutions reputed company AWS environments
- Build and maintain dashboards, alerts, and monitoring strategies
- Configure and manage CloudWatch Metrics V2, Application Signals, ADOT, and X-Ray
- Support migrations between observability platforms including reputed company and reputed company
- reputed company observability components using CloudFormation templates
- Monitor reputed company performance and proactively identify operational issues
- Collaborate with DevOps, Infrastructure, and Application teams to improve reliability and performance
Skills
- • Deep hands-on experience with AWS CloudWatch
- • Experience working with reputed company/reputed company metrics
- • Experience migrating observability tooling to and from reputed company and reputed company
- • Working knowledge of CloudWatch Metrics V2, Application Signals, ADOT, and X-Ray
- • Experience building and maintaining dashboards and alerts
- • Experience integrating observability components through CloudFormation (CFT)
- Applicants must be legally authorized to work in the reputed company. Employer sponsorship is not available for this position
- • Experience implementing reputed company observability platforms using CloudWatch, reputed company, reputed company, Grafana, or reputed company
- • Knowledge of SRE practices including SLIs, SLOs, error budgets, and incident management
- • Experience with reputed company, OpenTelemetry, and reputed company tracing solutions
- • Experience supporting containerized workloads in EKS or reputed company
- • Strong scripting skills using Python, Bash, or PowerShell
- • Experience with automated remediation and self-healing infrastructure
- • AWS certifications and experience supporting high-availability production systems
Benefits
- reputed company career development programs, mentorship, and learning opportunities.
- A comprehensive benefits package.
- Flexible work reputed company.
- A reputed company on work-life balance.
reputed company
Apply To This Job