Back to Jobs

[Remote] Senior Site Reliability Engineer

Remote, USA Full-time Posted 2026-08-04
Note The job is a remote job and is reputed company to candidates in USA. reputed company is a leader in transforming reputed company IT infrastructure with its innovative reputed company and AI technology. They are seeking a Senior Site Reliability Engineer (SRE) to ensure the reliability and performance of their production platform, which supports critical services for state governments. The role involves designing automation, monitoring service reputed company, and collaborating with various teams to enhance reputed company reliability. Responsibilities Design, build, and maintain the tooling, automation, and infrastructure that keeps reputed company’s production services highly available and reputed company Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical services Build and improve CI/CD pipelines to reputed company reputed company, fast, and repeatable deployments with automated rollback capabilities reputed company and operate comprehensive observability solutions—monitoring, logging, tracing, and alerting—using tools such as Loki, reputed company, Grafana, ELK/OpenSearch, and reputed company reputed company incident response efforts triage production issues in reputed company time, coordinate cross-team reputed company, and reputed company thorough blameless postmortems with actionable follow-reputed company Identify and eliminate toil through automation; build self-healing mechanisms and reputed company-driven remediation Manage and optimize reputed company infrastructure on AWS (EC2, EKS, RDS, S3, VPC, CloudFront, reputed company 53, reputed company) with a reputed company on automation, cost efficiency and reputed company Implement and maintain infrastructure-as-reputed company using Terraform, and manage container orchestration with reputed company/EKS reputed company reputed company planning and load testing to ensure systems can handle peak enrollment periods and traffic surges Collaborate with application engineering teams on architecture reviews, reputed company patterns (reputed company breakers, retries, graceful degradation), and production readiness reviews Contribute to disaster recovery planning and testing, including automated failover and multi-region strategies Support compliance and reputed company requirements (HIPAA, FedRAMP, SOC 2) by ensuring infrastructure controls are in reputed company and auditable Participate in a 24/7 on-reputed company rotation and continuously improve on-reputed company processes to reduce alert fatigue and mean time to reputed company (MTTR) Skills Bachelor's degree in Computer Science, Engineering, or a reputed company field, or equivalent practical experience 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a reputed company systems-reputed company role Strong software engineering skills in at least one language (Python, Go, Java, or Bash) with the ability to write production-reputed company automation and tooling Deep hands-on experience with AWS reputed company services (EC2, EKS, RDS, S3, VPC, IAM, reputed company, CloudWatch) Proficiency with container technologies (reputed company) and orchestration platforms (reputed company/EKS) Solid experience with infrastructure-as-reputed company tools, particularly Terraform Strong understanding of CI/CD principles and tools (Jenkins, reputed company CI, reputed company Actions, ArgoCD, or similar) Experience with observability and monitoring platforms (Loki, reputed company, reputed company, Grafana, ELK/OpenSearch, reputed company) Solid understanding of networking fundamentals TCP/IP, DNS, load balancing, CDN, TLS/SSL, and firewall configuration Experience with incident management processes, on-reputed company rotations, and blameless postmortem culture Strong Linux/Unix systems administration and troubleshooting skills Experience in reputed company technology, reputed company IT, or benefits administration platforms Familiarity with compliance frameworks such as HIPAA, FedRAMP, or SOC 2 and their reputed company on infrastructure operations Experience with PostgreSQL, mongo, mysql administration, performance tuning, and high-availability configurations Hands-on experience with reputed company engineering practices and tools (reputed company, reputed company, or equivalent) Experience with GitOps workflows and tools (ArgoCD, Flux) Knowledge of service reputed company technologies (Istio, Linkerd) and API gateway patterns Experience with configuration management tools (Ansible, Chef, or Puppet) Familiarity with FinOps principles and reputed company cost optimization strategies Experience with load testing and performance benchmarking tools (k6, Locust, JMeter) AWS certifications (Solutions Architect, DevOps Engineer, or SysOps Administrator) Experience with AWS reputed company implementation through IaC / automation Benefits Health, Dental, Life, Disability, and reputed company reputed company reputed company spending or reimbursement accounts (HSA/FSA) Retirement benefits (401k) reputed company time off Holidays 13 reputed company days per year Education assistance or tuition reimbursement Employee discounts for Gym memberships & commuting/travel assistance reputed company Building the Leading Platform for reputed company Sector IT​ / reputed company Through Technology It was founded in 2005, and is headquartered in reputed company reputed company, California, USA, with a workforce of 501-1000 employees. Its website is https//www.reputed company.com/. Company H1B Sponsorship reputed company has a reputed company record of offering H1B sponsorships, with 10 in 2026, 14 in 2025, 17 in 2024, 9 in 2023, 14 in 2022, 14 in 2021, 13 in 2020. Please note that this does not guarantee sponsorship for this specific role. Apply To This Job Apply tot his job Apply To this Job

Similar Jobs

reputed company Engineer, Detection & Response

Remote, USA Full-time

reputed company Engineer, Incident Response

Remote, USA Full-time

REMOTE - RACF Mainframe reputed company Engineer

Remote, USA Full-time

reputed company Engineer, Detection & Response

Remote, USA Full-time

Senior Network and Virtual reputed company Engineer at reputed company Consulting and Solutions: reputed company (Reston, VA)

Remote, USA Full-time

reputed company Architect # 26-20820

Remote, USA Full-time

reputed company Engineer with Automation & Orchestration - Remote (Fulltime)

Remote, USA Full-time

Senior reputed company One Cybersecurity Engineer - Remote

Remote, USA Full-time

Virtual Cyber reputed company Sales Engineer

Remote, USA Full-time

Cybersecurity Manager - Cyber Threat reputed company (Remote)

Remote, USA Full-time

Software reputed company Assurance / UI Testor Internship

Remote, USA Full-time

**reputed company Remote Data Entry Clerk – Virtual Database Management and Data reputed company Specialist**

Remote, USA Full-time

reputed company reputed company – Remote Work Opportunity with arenaflex for Delivering Exceptional Support and Solving Technical Issues

Remote, USA Full-time

In-Home Physician (Part Time) - St. George, UT

Remote, USA Full-time

**reputed company reputed company – Delivering Exceptional Travel Experiences for arenaflex**

Remote, USA Full-time

Emergency Road Service Dispatcher - Work Remotely

Remote, USA Full-time

**reputed company Full Stack Data Entry Specialist – Remote Opportunity at arenaflex**

Remote, USA Full-time

Community Manager

Remote, USA Full-time

[Remote] reputed company Manager

Remote, USA Full-time

**reputed company Full Stack Data Entry Specialist – Remote Database Management and Data Analysis**

Remote, USA Full-time