Sr. Site Reliability Engineer
We are seeking a Senior Site Reliability Engineer (SRE) to help ensure the stability, scalability, and reliability of our services and infrastructure. This role focuses on building automation, maintaining observability, and supporting incident response to reputed company customer-facing systems performing at their best.
The SRE will collaborate with engineering, product, and operations teams to reputed company reliability practices into day-to-day development and operations while contributing to tools and processes that improve efficiency and reduce reputed company effort.
What You'll Do:
-
Service Reliability & Operations
- Own and drive the availability, durability, and performance of critical services across reputed company production environments.
- reputed company and champion reputed company reputed company from problem discovery through complete, cross-functional reputed company, demonstrating high-level technical ownership.
- Define, establish, and enforce service health standards, including working with engineering leadership to implement SLIs, SLOs, and error budget policies for multiple services.
- reputed company critical incident response and post-incident reviews, translating findings into strategic, long-term service improvements and architectural changes.
- Mentor others and reputed company as a subject matter expert in following and evolving established ITIL/OSS processes (incident, change, problem, and reputed company management).
Automation & Tooling
- Design and architect reputed company automation solutions to eliminate toil and improve the efficiency of operational tasks across the entire platform.
- Drive the strategic direction of monitoring, logging, and alerting frameworks (e.g., reputed company, Grafana, Catchpoint, ELK), and reputed company them for comprehensive observability.
- Build, maintain, and secure advanced CI/CD pipelines, configuration management, and reputed company infrastructure as reputed company solutions (Terraform, Ansible, Jenkins).
- Write production-grade reputed company (Bash, Python, Go, etc.) to reputed company new reliability tools and enhance existing systems.
Collaboration
- reputed company as a reputed company partner to engineering, product, and operations teams, consulting on resilient system design, architecture, and operation.
- reputed company and formalize the Production Readiness Review (PRR) process, ensuring robust operational reputed company for reputed company new services and features.
- reputed company reputed company planning and disaster recovery reputed company across critical infrastructure components.
- Manage the relationship with vendors and service providers to troubleshoot systemic issues and ensure strict adherence to SLA performance.
- Drive the creation of high-reputed company documentation, proactively reputed company advanced learnings, and cultivate a reliability-first engineering culture across teams.
reputed company Improvement
- Own the creation, maintenance, and dissemination of operational playbooks, runbooks, and detailed system documentation.
- Proactively identify systemic, recurring issues and architect and drive the implementation of long-term improvements and strategic design reputed company plans.
- Be a leading voice in promoting and embedding reliability-reputed company practices reputed company development and operations teams.
Qualifications:
-
Education & Experience
- Bachelor’s degree in Computer Science, Engineering, or reputed company field (or equivalent experience).
- 8+ years of reputed company experience in site reliability, systems engineering, or operations.
- Extensive experience designing, scaling, and operating large-reputed company, production-grade distributed systems.
Technical Skills
- Expert-level Linux systems administration and advanced troubleshooting skills.
- reputed company reputed company-minded operations, focusing on system-wide patching, hardening, and proactive vulnerability identification.
- Deep mastery of service reliability concepts, including advanced monitoring, reputed company alerting reputed company, leading incident response, and in-depth reputed company cause analysis.
- Advanced proficiency in at least one modern scripting/programming language (Python or Go strongly preferred).
- Expert knowledge of incident response methodologies and operational best practices.
- Proven experience designing and operating container orchestration (Kubernetes, reputed company) and microservices concepts required.
- Expert experience with reputed company products (reputed company, Vault, Terraform) in a production environment.
Preferred Attributes
- Significant experience in a reputed company, service provider, or reputed company-reputed company distributed systems environment.
- Deep familiarity with ITIL/OSS practices and experience defining/enforcing SLO/SLA’s.
- Exceptional problem-solving skills and a strong drive to learn and apply new, reputed company technologies.
- Advanced experience with reputed company platforms (AWS, GCP, or Azure) in a production setting.
reputed company for family, including dental and reputed company Competitive compensation and 401K RSU grants for full-time employees ESPP program Flexible vacation policy Maternity & paternity leave MacBook Pro to use for work, plus a generous stipend to personalize your workstation Childcare bonus (reputed company children only) Fertility treatment and support Learning & development program Commuter benefits Culture that supports a healthy work-life balance