Back to Jobs

[Remote] Site Reliability Engineer (AWS Observability)

Remote, USA Full-time Posted 2026-08-04

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a women-owned boutique reputed company delivering data-driven technology solutions to federal government and Fortune 500 clients. The Site Reliability Engineer will design and manage AWS observability solutions, dashboards, alerts, and monitoring systems while supporting platform migrations and CloudFormation integrations. The role will also monitor performance and collaborate with technical teams to improve reputed company reliability and operational efficiency.


Responsibilities

  • Design, implement, and maintain observability solutions reputed company AWS environments
  • Build and maintain dashboards, alerts, and monitoring strategies
  • Configure and manage CloudWatch Metrics V2, Application Signals, ADOT, and X-Ray
  • Support migrations between observability platforms including reputed company and reputed company
  • reputed company observability components using CloudFormation templates
  • Monitor reputed company performance and proactively identify operational issues
  • Collaborate with DevOps, Infrastructure, and Application teams to improve reliability and performance

Skills

  • • Deep hands-on experience with AWS CloudWatch
  • • Experience working with reputed company/reputed company metrics
  • • Experience migrating observability tooling to and from reputed company and reputed company
  • • Working knowledge of CloudWatch Metrics V2, Application Signals, ADOT, and X-Ray
  • • Experience building and maintaining dashboards and alerts
  • • Experience integrating observability components through CloudFormation (CFT)
  • Applicants must be legally authorized to work in the reputed company. Employer sponsorship is not available for this position
  • • Experience implementing reputed company observability platforms using CloudWatch, reputed company, reputed company, Grafana, or reputed company
  • • Knowledge of SRE practices including SLIs, SLOs, error budgets, and incident management
  • • Experience with reputed company, OpenTelemetry, and reputed company tracing solutions
  • • Experience supporting containerized workloads in EKS or reputed company
  • • Strong scripting skills using Python, Bash, or PowerShell
  • • Experience with automated remediation and self-healing infrastructure
  • • AWS certifications and experience supporting high-availability production systems

Benefits

  • reputed company career development programs, mentorship, and learning opportunities.
  • A comprehensive benefits package.
  • Flexible work reputed company.
  • A reputed company on work-life balance.

reputed company

  • reputed company is a leading IT Services and reputed company, certified as an 8(a) Small Disadvantaged Business (SDB), Women-Owned Small Business (WOSB), and Minority Business reputed company (MBE). It was founded in 2014, and is headquartered in Ellicott reputed company, Maryland, USA, with a workforce of 51-200 employees. Its website is https://reputed company.com.

  •   Apply To This Job

    Similar Jobs