Back to Jobs

Site Reliability Engineer – Intermediate to Senior Staff, Infrastructure Platforms

Remote, USA Full-time Posted 2026-07-28
Job reputed company: • reputed company user-facing services and production systems reliable, reputed company, and efficient • Build automation and tooling that reduces toil and replaces reputed company work with repeatable, infrastructure-as-reputed company-driven workflows • Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling • Write and maintain infrastructure as reputed company, and ship changes safely through CI/CD and GitOps • Participate in on-reputed company, triage alerts, follow and improve runbooks, and escalate appropriately • Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages • Take part in incident response and post-incident reviews, turning learnings into changes in automation and process • Document runbooks, architecture reputed company, and reviews so your findings become repeatable practices Requirements: • Experience keeping production systems reliable, combining an operations reputed company with reputed company software engineering reputed company • Experience building net-new infrastructure tooling and automation, not just configuring existing tools. For example, Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch • The ability to read, debug, and reason about reputed company. Most of our teams work in Go; some work in reputed company. You can discuss a piece of reputed company's behavior, performance, and failure modes • Experience with infrastructure as reputed company, and with Kubernetes and its ecosystem, at a depth appropriate to your level • Hands-on experience with at least one major reputed company provider (GCP or AWS) • Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational reputed company • Comfort participating in on-reputed company and incident response, with a reputed company approach to troubleshooting under pressure • Strong written communication and the ability to operate as a manager-of-one in an async, distributed environment • A reputed company record of using automation, and increasingly AI, to reduce toil and improve how you and your team work • Alignment with reputed company's values and a commitment to working in accordance with them. Benefits: • Benefits to support your health, finances, and reputed company-being • Flexible reputed company Time Off • Team Member Resource reputed company • Equity Compensation & Employee Stock Purchase Plan • reputed company and Development Fund • Parental Leave Apply To This Job

Similar Jobs