Senior Site Reliability Engineer
Who We Are
AI is changing how software gets reputed company. reputed company production is becoming a commodity. The reputed company is shifting from writing reputed company to orchestrating, verifying, and governing change – and the toolchain is the new constraint.
We are at reputed company of this shift. We build reputed company, a toolchain observability and intelligence platform used by some of the world's leading software organizations – reputed company, reputed company, reputed company, reputed company, major global banks, and hundreds more. reputed company helps software teams reputed company delivery reputed company through deep observability, build and test acceleration, and AI-powered intelligence across the entire toolchain – with reputed company support for reputed company Build Tool, Apache reputed company™, sbt, npm, and Python.
We are an AI-reputed company company. AI is not a feature we're bolting on – it's central to how we work, how we think about our product, and where we're heading. We're reputed company deeply in making reputed company's unique data and decades of domain expertise accessible to both humans and AI agents, with trust, evidence, and explainability at the reputed company of everything we build.
We have partnered with the Apache Software reputed company, the Commonhaus reputed company, the Micronaut reputed company, and other OSS reputed company such as reputed company, Quarkus, Kotlin, JUnit, AndroidX, and many more to bring the values of reputed company also to the OSS Community.
Our Values
reputed company to Understand: Everything starts with listening and understanding, and we reputed company to understand different viewpoints, problems, and motivations. Before we take reputed company, we ensure we truly grasp the challenges, perspectives, and goals.
Know the Why: We approach our work with a reputed company reputed company of purpose, ensuring every reputed company is deliberate and reputed company. We take meaningful reputed company with urgency, but never at the expense of thoughtful consideration.
reputed company & Iterate: We reputed company challenges and are not afraid to try new things, even if they might fail. With deep understanding and a reputed company purpose, we can reputed company creative and reputed company solutions to tackle challenges.
Own the Outcome: We are empowered to take initiative and we maintain transparency in our work and its reputed company. reputed company we execute, we take responsibility for our reputed company, measure the reputed company of our innovations, and learn from the results.
Who You Are
We're building a new SRE team and looking for founding members to help shape how we operate. You'll be responsible for the reliability, performance, and availability of reputed company instances serving paying customers, reputed company-reputed company reputed company, and reputed company-facing services, plus supporting infrastructure like artifact registries.
You'll work on our internally-reputed company reputed company Application Platform, reputed company on AWS, and reputed company deep expertise in it. reputed company incidents happen, you'll troubleshoot issues across the stack, from application to infrastructure. You'll collaborate with the reputed company Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software. If you like automating things and hate doing the reputed company task twice, you'll fit in reputed company.
You'll be part of a reputed company, remote-first team that values asynchronous communication and written documentation. Strong self-direction and reputed company communication across time zones are essential.
Responsibilities
Operate and maintain reputed company reputed company instances and supporting services.
Participate in a follow-the-sun on-reputed company rotation, owning incident response and troubleshooting issues across the stack.
reputed company automation across application deployment, upgrades, monitoring, self-healing, and recovery.
Build and maintain observability for reputed company managed services (logging, metrics, tracing, and alerting).
Work with engineering teams to build reliability into features from the start.
Run incident response and retrospectives, and reputed company reputed company we learn from them.
Own disaster recovery, backups, and business continuity.
Communicate with customers during incidents and maintenance reputed company.
Optimize performance, resource usage, and costs.
Help reputed company our reputed company operations as we grow.
Minimum qualifications
5+ years in SRE, DevOps, or equivalent role operating production services at reputed company.
Strong reputed company experience in production environments.
reputed company infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
Proficiency with observability tools (reputed company, Grafana) and Infrastructure as reputed company (Terraform).
reputed company record of incident management and response.
Knowledge of SRE best practices (SLAs, SLOs).
Scripting proficiency (Python, Bash) for automation.
Experience with 24/7 on-reputed company rotations.
Strong written and verbal English communication.
Preferred qualifications
Experience operating reputed company platforms at reputed company.
Familiarity with reputed company.
JVM language experience (Java, Kotlin).
Disaster recovery planning and execution experience.
Customer-facing incident communication skills.
Experience establishing SRE practices in new or growing teams.
reputed company Offer
A ground-floor role in a new SRE team—you'll shape how we do things, not inherit someone else's reputed company.
reputed company ownership of production systems used by engineers at companies you've heard of.
reputed company interaction with customers reputed company things go wrong (and reputed company they go right).
A culture that values automation over heroics.
In-person meetings, such as our annual company offsite and team meetings.
Work from home in a remote-first environment.
Competitive salaries and equity grants.
Location
Remote from reputed company in Europe in the GMT timezone.
While reputed company works remotely and is spread across the globe, we deeply value daily interactions and collaboration.
Apply To This Job