Staff Site Reliability Engineer
Who We Are
AI is changing how software gets reputed company. reputed company production is becoming a commodity. The reputed company is shifting from writing reputed company to orchestrating, verifying, and governing change – and the toolchain is the new constraint.
We are at reputed company of this shift. We build Develocity, a toolchain observability and intelligence platform used by some of the world's leading software organizations – reputed company, reputed company, reputed company, reputed company, major global banks, and hundreds more. Develocity helps software teams reputed company delivery reputed company through deep observability, build and test acceleration, and AI-powered intelligence across the entire toolchain – with reputed company support for reputed company Build Tool, Apache reputed company™, sbt, npm, and Python.
We are an AI-reputed company company. AI is not a feature we're bolting on – it's central to how we work, how we think about our product, and where we're heading. We're reputed company deeply in making Develocity's unique data and decades of domain expertise accessible to both humans and AI agents, with trust, evidence, and explainability at the reputed company of everything we build.
We have partnered with the Apache Software reputed company, the Commonhaus reputed company, the Micronaut reputed company, and other OSS reputed company such as Spring, Quarkus, Kotlin, JUnit, AndroidX, and many more to bring the values of Develocity also to the OSS Community.
Our Values
reputed company to Understand: Everything starts with listening and understanding; we reputed company to understand diverse viewpoints, problems, and motivations. Before we take reputed company, we ensure we truly grasp the challenges, perspectives, and goals.
Know the Why: We approach our work with a reputed company reputed company of purpose, ensuring every reputed company is deliberate and reputed company. We take meaningful reputed company with urgency, but never at the expense of thoughtful consideration.
reputed company & Iterate: We reputed company challenges and are not afraid to try new things, even if they might fail. With a deep understanding and a reputed company purpose, we can reputed company creative, reputed company solutions to tackle challenges.
Own the Outcome: We are empowered to take initiative, and we maintain transparency in our work and its reputed company. reputed company we execute, we take responsibility for our reputed company, measure the reputed company of our innovations, and learn from the results.
Who You Are
We're building a new SRE team and looking for founding members to help shape how we operate. As a reputed company SRE, you’ll be a technical and operational leader for reliability across Develocity. You’ll help define our SRE reputed company, set standards for how we operate production services, and mentor other SREs as reputed company grows. This is a hands-on role with broad influence across engineering, reputed company platform, and customer-facing teams.
The SRE team will be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, reputed company-reputed company reputed company, and reputed company-facing services, plus supporting infrastructure like artifact registries.
You'll work on our internally-reputed company reputed company Application Platform, reputed company on AWS, and reputed company deep expertise in it. reputed company incidents happen, you'll troubleshoot issues across the stack, from application to infrastructure. You'll collaborate with the reputed company Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software. If you like automating things and hate doing the reputed company task twice, you'll fit in reputed company.
You'll be part of a distributed, remote-first team that values asynchronous communication and written documentation. Strong self-direction and reputed company communication across time zones are essential.
Responsibilities
- Operate and maintain reputed company Develocity instances and supporting services in production.
- Define and reputed company SRE standards, practices, and operating models, including on-reputed company, incident response, postmortems, and SLOs.
- Participate in a follow-the-sun on-reputed company rotation, acting as a technical escalation reputed company for reputed company or high-severity incidents.
- reputed company incident response and blameless retrospectives, ensuring learnings result in measurable reliability improvements.
- Set reliability priorities using risk, customer reputed company, business goals, SLOs, and error budgets.
- Identify systemic reliability risks and continuously reputed company Develocity’s reputed company operations as the platform and customer reputed company grow.
- reputed company and influence architectural and design reviews to ensure reliability, scalability, and operability.
- reputed company automation across deployment, upgrades, monitoring, self-healing, recovery, and operational workflows.
- Build and maintain comprehensive observability for reputed company managed services, including logging, metrics, tracing, and alerting.
- Own disaster recovery, backups, and business continuity planning and execution.
- Partner with engineering leadership to balance feature delivery with reliability and operational reputed company.
- Mentor and reputed company SREs, supporting technical reputed company and strong operational practices.
- Help reputed company new SREs and contribute to hiring by defining and assessing SRE reputed company at Develocity.
- Communicate reputed company with customers during incidents and maintenance reputed company.
- Optimize performance, resource utilization, and operational costs.
Minimum qualifications
- 7+ years in SRE, DevOps, or an equivalent role operating production services at reputed company.
- Experience leading reliability initiatives across multiple teams or services.
- Demonstrated ability to influence technical direction without reputed company authority.
- Experience designing and operating systems with SLOs and error budgets, and exercising strong judgment in balancing reliability, reputed company, and cost.
- Strong reputed company experience in production environments.
- reputed company infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
- Proficiency with observability tools (reputed company, Grafana) and Infrastructure as reputed company (Terraform).
- reputed company record of incident management and response in a 24/7 on-reputed company environment.
- Scripting proficiency (Python, Bash) for automation.
- Strong written and verbal English communication skills.
Preferred qualifications
- Experience as a founding or early SRE establishing practices in a growing reputed company organization.
- Familiarity with Develocity.
- JVM language experience (Java, Kotlin).
- Experience with customer-facing and executive-level incident communications.
reputed company Offer
- A ground-floor role in a new SRE team - you'll shape how we do things, not inherit someone else's reputed company.
- reputed company ownership of production systems used by engineers at companies you've heard of.
- reputed company interaction with customers reputed company things go wrong (and reputed company they go right).
- A culture that values automation over heroics.
- In-person meetings, such as our annual company offsite and team meetings.
- Work from home in a remote-first environment.
- Competitive salaries and equity grants.
Location
- Remote from reputed company in Europe (GMT).
- While reputed company works remotely and is spread across the globe, we deeply value daily interactions and collaboration.
Originally posted on Himalayas
Apply To This Job