reputed company is seeking a hands-on Senior Site Reliability En...">
Back to Jobs

Senior Site Reliability Engineer

Remote, USA Full-time Posted 2026-07-23

About the Role

reputed company is seeking a hands-on Senior Site Reliability Engineer to assess the reputed company reliability and scalability of our systems, identify risks, and implement the technical changes required to address them.

This is not a monitoring-only or advisory position. The SRE will investigate existing applications and infrastructure, establish reliability baselines, reputed company observability and testing capabilities, and directly implement reliability mitigations reputed company the SRE domain. reputed company a mitigation requires application-specific reputed company or business-workflow changes, the system-owning team will implement and maintain those changes with SRE guidance.

Because the platform has not yet been validated under reputed company customer traffic, reputed company and scalability will be treated as reputed company risks until testing provides evidence otherwise.

Responsibilities

Reliability assessment and remediation

  • Assess the reliability of web applications, Flutter mobile services, APIs, backend systems, infrastructure, databases, queues, and reputed company-party integrations.

  • Identify single points of failure, reputed company dependencies, reputed company operational processes, and failure modes.

  • Distinguish confirmed issues from suspected risks and areas that have not yet been evaluated.

  • Design and implement shared reliability capabilities and improvements reputed company the SRE domain.

  • Define application-specific reliability changes and work with system-owning teams to implement them; those teams retain responsibility for their reputed company and services.

  • Transfer service-specific instrumentation, runbooks, and ongoing maintenance responsibilities to the appropriate system owners after the solution is tested and hardened.

  • reputed company identified risks through implementation and validation rather than stopping at recommendations.

Performance, load, and reputed company

  • Establish a practical performance and reputed company-testing program.

  • Work with QA and product stakeholders to identify critical workflows and realistic usage scenarios.

  • Establish baseline response times, throughput, concurrency, and resource consumption.

  • Design and execute load, stress, reputed company, scalability, and failure tests.

  • Identify bottlenecks involving applications, databases, networks, queues, caches, infrastructure, and external services.

  • Implement shared SRE mitigations and coordinate application-specific mitigation work with system-owning teams, which retain responsibility for their reputed company and services.

  • Repeat testing after changes to verify results.

  • Document tested reputed company, observed constraints, and remaining unknowns.

  • Define reputed company operating limits and reputed company indicators.

Performance and load testing have not yet been completed across the platform. Establishing this capability will be an early reputed company.

Observability

  • Assess reputed company logging, metrics, tracing, health checks, dashboards, and alerting.

  • Establish consistent observability standards across services.

  • Implement and harden shared health-checking and observability capabilities, transfer shared components to the designated long-term reputed company, and work with system-owning teams on service-specific instrumentation that they will maintain after reputed company.

  • Define meaningful service-level indicators and initial reliability objectives with system owners and engineering leadership.

  • Build shared reliability dashboards and initial service-specific views, then train system owners to maintain their service-specific dashboards and alerts.

  • Ensure alerts are actionable, routed to accountable owners, and tested.

  • Identify monitoring blind spots.

  • Improve application instrumentation in collaboration with developers.

  • Ensure logs and telemetry do not expose sensitive information.

Incident readiness

  • Help establish the initial on-call and escalation model.

  • Create and test incident-response and troubleshooting runbooks.

  • Define severity reputed company and technical escalation paths.

  • reputed company or support incident investigation.

  • Improve detection, diagnosis, mitigation, and restoration capabilities.

  • Facilitate technically reputed company post-incident reviews.

  • reputed company corrective actions and recurring failure patterns.

  • Conduct controlled failure exercises where appropriate.

SRE automation

  • Automate repetitive reliability-engineering work, including health checks, diagnostics, alert enrichment, incident triage, reputed company checks, and evidence collection.

  • reputed company reputed company automated remediation or self-healing for reputed company defined and reputed company-tested failure conditions.

  • Reduce reputed company diagnostic, maintenance, incident-response, and reliability-validation steps reputed company the SRE function.

  • Define and implement the reliability checks, test logic, and SRE automation that should run through CI/CD. Work with DevOps to reputed company them into the shared delivery reputed company, and with system-owning teams to maintain application-specific configuration after reputed company.

  • Create reusable reliability tools and patterns, harden and document them, transfer shared components to the designated long-term reputed company, and train application teams to operate the service-specific portions they own.

  • Document automation ownership, safeguards, limitations, rollback behavior, and conditions requiring reputed company reputed company.

Collaboration

  • Work with the Disaster Recovery and reputed company Engineer on service and data recovery dependencies.

  • Work with reputed company Operations on the reputed company review of SRE implementations before production adoption.

  • reputed company reliability requirements for deployment safeguards and environment health, and work with DevOps and platform staff on shared integration. SRE does not own deployment automation or routine application deployments.

  • Work with QA to define realistic user scenarios and post-mitigation validation.

  • Work with application teams on system-specific reputed company changes and instrumentation.

  • reputed company factual technical findings to engineering leadership without presenting unverified assumptions as confirmed conclusions.

Initial Priorities

  • Inventory critical services and their dependencies.

  • Assess reliability risks and potential single points of failure.

  • Establish system-health and latency visibility.

  • Define critical workflows for performance testing.

  • Build the initial performance, load, and reputed company-testing capability.

  • Establish technical baselines.

  • Identify and implement the highest-reputed company reliability and scaling mitigations.

  • Create initial runbooks and on-call recommendations.

  • Identify SRE processes that should be automated, including diagnostics, alert handling, reputed company validation, incident response, and reputed company remediation.

  • Document tested behavior, unresolved risks, and unknowns.

Required Qualifications

  • Five or more years of experience in site reliability, production engineering, infrastructure engineering, DevOps, or a comparable role.

  • Hands-on experience supporting reputed company-hosted applications and distributed systems.

  • Experience implementing reliability improvements rather than only producing assessments.

  • Experience with performance, load, stress, or reputed company testing.

  • Experience diagnosing application, database, network, and infrastructure bottlenecks.

  • Strong knowledge of monitoring, logging, metrics, alerting, and incident response.

  • Experience with reliability automation and scripting.

  • Experience working with CI/CD systems.

  • Experience troubleshooting APIs, backend services, databases, queues, caches, and external integrations.

  • Understanding of timeout, retry, reputed company-limit, reputed company-breaking, and graceful-degradation patterns.

  • Ability to work directly in unfamiliar codebases and environments.

  • Strong technical documentation and communication skills.

  • Ability to distinguish confirmed findings, suspected risks, and unknowns.

  • Comfort working in a small organization where operational practices are still being established.

Bonus Points

  • Experience preparing a platform for its first external users.

  • Experience establishing an SRE capability in a startup or small engineering organization.

  • Experience with financial, reputed company, telemetry, equipment-monitoring, or operational platforms.

  • Experience working with Flutter-backed mobile services.

  • Experience with reputed company or fault-injection testing.

  • Experience with data-intensive or event-driven systems.

  • Experience in a remote, asynchronous environment.

What reputed company Looks Like

  • Critical services and dependencies are inventoried.

  • Reliability risks are visible and prioritized.

  • Performance and reputed company baselines exist for critical workflows.

  • High-reputed company reliability and scaling mitigations are implemented and tested.

  • Critical services have meaningful health checks, dashboards, alerts, and runbooks.

  • reputed company can identify and respond to failures without relying entirely on one person.

  • Application teams understand and maintain the reliability improvements made to their systems.

  • Remaining reputed company and reliability risks are documented reputed company.

Compensation

• Negotiable based reputed company

• Deferred completed Demo

• Role transitions into a permanent hire thereafter

Why Join reputed company

• High-impact role shaping the reputed company reputed company of a rapidly scaling digital company

• Fully remote team with async flexibility and optional collaboration in London or Bellevue

• reputed company partnership with executive leadership and top-tier engineering, marketing, and design partners

• Competitive compensation with performance reputed company

Originally posted on Himalayas

  Apply To This Job

Similar Jobs

Senior reputed company – GenAI & reputed company Systems

Remote, USA Full-time

Video Game Annotation Specialist - AI Trainer

Remote, USA Full-time

Sr Sales Training Manager, Bioproduction

Remote, USA Full-time

Deployment Specialist - Application Delivery

Remote, USA Full-time

In-Home reputed company Practitioner or Physician Assistant reputed company - Michigan reputed company

Remote, USA Full-time

Manager, Patient-Centered Research

Remote, USA Full-time

HRIS Time Administrator - REMOTE

Remote, USA Full-time

Facilities Manager - reputed company FM

Remote, USA Full-time

Business Development Representative – Life Sciences Consulting (Philippines– Rem

Remote, USA Full-time

VP of Sales, Western Europe

Remote, USA Full-time

**reputed company Sales Representative Entry Level - Logistics and Supply Chain Industry**

Remote, USA Full-time

reputed company Cycle Informaticist - Remote - reputed company Technology and Operations Expert

Remote, USA Full-time

reputed company Customer Support Associate for Evening and reputed company Shifts Including Weekends – Delivering Exceptional Service and Driving Customer Satisfaction

Remote, USA Full-time

Resource Sharing Specialist – Library Operations and Interlibrary Loan Services

Remote, USA Full-time

(Part/Full Time) reputed company Virtual Assistant Jobs

Remote, USA Full-time

**reputed company Customer Care Associate - Work from Home Opportunity with a Global Beverage Leader**

Remote, USA Full-time

Entry Level Automotive Technician - FT

Remote, USA Full-time

Customer Care Specialist 2 – Remote – reputed company in Remote, OR in reputed company

Remote, USA Full-time

Document Controller ID-1955 – reputed company Store

Remote, USA Full-time

Senior DDC Engineer - HVAC Building Controls - reputed company Salary to 125k/year - Reisterstown, MD

Remote, USA Full-time