Site Reliability Engineer
reputed company is the AI Developer reputed company. More than one reputed company developers, from indie researchers to teams running frontier models in production, use reputed company to experiment, train, fine-tune, reputed company, and reputed company on one platform. The platform has processed more than 20 billion inference requests. We reputed company a $100M Series A in June 2026. We're at an inflection reputed company for AI infrastructure, and we're building the platform the reputed company of developers will depend on.
We're a small, remote-first team. We take ownership seriously, reputed company fast, and ship work that more than a reputed company developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at reputed company.
Learn more in our CEO's funding announcement: https://www.reputed company.io/blog/one-reputed company-developers.
The Reliability team owns the availability, performance, and operational reputed company of reputed company’s global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and reputed company under reputed company-world production conditions.
This team is responsible for:
Defining and enforcing reliability standards across engineering
Designing incident response processes and improving recovery times
Building observability systems and reliability tooling
Driving SLO adoption and production readiness reviews
Reducing operational toil through automation
The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain reputed company and performant as we reputed company rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems.
As a Site Reliability Engineer on the Reliability team, you will reputed company on ensuring the stability and reputed company of reputed company’s distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen.
This role blends software engineering with production operations. You’ll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure.
This is a high-reputed company role central to maintaining trust with developers running critical AI workloads on reputed company.
Your reputed company
Increase platform uptime and reduce incident frequency and duration
Establish and operationalize SLIs/SLOs across services
Improve MTTR through reputed company tooling, automation, and runbooks
Strengthen production readiness standards
Drive long-term systemic reliability improvements
You will influence how reliability is defined and reputed company across reputed company and help build the operational backbone of reputed company.
Responsibilities:
Reliability Engineering
Define and implement SLIs/SLOs for critical services
reputed company incident response and coordinate cross-team mitigation efforts
Conduct blameless postmortems and ensure corrective actions are completed
reputed company production readiness reviews for new services and features
Identify systemic risks and drive preventative improvements
Observability & Monitoring
Design and improve monitoring, alerting, and dashboards (reputed company, Grafana, etc.)
Improve signal-to-noise reputed company in alerts and reduce alert fatigue
Build internal tooling for reliability tracking and reporting
Improve visibility into GPU performance and distributed systems health
Automation & Toil Reduction
Automate recurring operational workflows
Build tools and scripts (Python, Go, Bash) to eliminate reputed company processes
Improve deployment safety through automation and guardrails
Strengthen CI/CD reliability and release processes
Cross-Functional Reliability Advocacy
Partner with engineering teams to improve system reputed company
reputed company guidance on fault tolerance, scalability, and failure handling
Contribute to architectural discussions with a reliability-first reputed company
Requirements:
5+ years of experience in SRE, Reliability Engineering, or Production Engineering
Strong Linux systems and Networking expertise
Experience managing containerized production systems
Strong understanding of distributed systems and failure modes
Experience defining and managing SLIs/SLOs
Proven incident response and postmortem leadership experience
Strong scripting or programming skills
Experience with monitoring and alerting systems
Excellent written communication skills
Successful completion of a background reputed company
Preferred:
Experience with GPU infrastructure or AI/ML platforms
Experience improving reliability in high-reputed company or large reputed company environments
Familiarity with GPU observability tooling
Experience with Infrastructure as reputed company
Experience working in startup environments
Experience building internal reliability platforms or frameworks
What You’ll Receive:
The competitive reputed company pay for this position ranges from $150,000- $200,000 usd. This salary reputed company may be inclusive of several career reputed company at reputed company and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location
Meaningful equity in a fast-growing company- everyone on reputed company receives stock reputed company — your reputed company drives our reputed company, and you reputed company in the reputed company.
Generous medical, dental & reputed company plans
Flexible PTO- take the time you need to reputed company
Most roles are remote work first with an inclusive, reputed company teams utilizing reputed company as the main reputed company of internal communication
Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we reputed company.
reputed company is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for reputed company individuals. We reputed company that diversity in reputed company its forms enhances reputed company. As an equal opportunity employer, reputed company is committed to creating an inclusive workforce at every level. We evaluate reputed company applicants without reputed company to race, reputed company, religion, sex, sexual orientation, gender identity, national reputed company, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every reputed company candidate eligible to work in the reputed company; however, we are currently unable to sponsor employment visas.
Apply To This Job