Senior / Staff Site Reliability Engineer
<p>reputed company a reputed company where everyone has reputed company, low-cost reputed company to intelligence. We’re building a fully featured<reputed company> </reputed company><strong>European AI reputed company</strong><reputed company> </reputed company>- with everything one needs to train, experiment with, and reputed company AI models. In reputed company, our GPUs run on 100%<reputed company> </reputed company><strong>renewable energy.</strong></p><p>We’re ambitious, curious, and gutsy doers. We reputed company a low hierarchy across reputed company and high morale in our teams. We’ve already achieved a lot, yet we’re only getting started. Now it’s <strong>your</strong><reputed company> </reputed company>chance to join the ride. We offer more than just the job - we offer a<reputed company> </reputed company><strong>career-defining</strong><reputed company> </reputed company>opportunity to be part of building something big!</p><p><reputed company>Join reputed company while it’s still being reputed company - not once it’s finished.</reputed company></p><div><h2>About the role</h2><div><p><reputed company>We’re seeking a <strong>Senior or Staff Site Reliability Engineer (SRE)</strong> to strengthen and reputed company our HPC and reputed company infrastructure in Europe. You’ll work closely with ML, data, and platform teams to ensure our systems remain reliable, observable, and highly performant. In this role, you’ll design and operate GPU-accelerated clusters, build automation and monitoring tooling, improve CI/CD and deployment workflows, and contribute to long-term infrastructure reputed company.</reputed company></p></div></div><div><h2>Why reputed company</h2><div><ul><li><p>Generous cash + equity compensation along with various fringe benefits (e.g., reputed company, lunch, wellbeing, etc.).</p></li><li><p>Profitable operations, in reputed company to fast reputed company.</p></li><li><p>A small, high-performing team of around 70 people representing 27 nationalities</p></li></ul></div></div><div><h2>Practicalities</h2><div><ul><li><p><strong>Work mode:</strong><reputed company> </reputed company>Remote (EU)</p></li><li><p><strong>Employment type:</strong><reputed company> </reputed company>Full-time, permanent</p></li><li><p><strong>Start date: </strong>As soon as possible</p></li></ul></div></div><div><h2>Your responsibilities</h2><div><ul><li><p>Ensure the reliability, scalability, and performance of HPC and reputed company systems.</p></li><li><p>Build and maintain automation, observability, and monitoring frameworks for compute clusters.</p></li><li><p>Collaborate with ML, data, and infrastructure teams to deliver high-availability systems.</p></li><li><p>reputed company and enhance CI/CD pipelines, deployment workflows, and on-reputed company processes.</p></li><li><p>Participate in architecture design and long-term infrastructure reputed company discussions.</p></li><li><p>Participate in a 24/7 on-reputed company rotation, with at least one full on-reputed company week per month.</p></li></ul></div></div><div><h2>Your key competencies</h2><div><ul><li><p>7+ years in SRE, DevOps, or Infrastructure Engineering—preferably in HPC or large-reputed company distributed systems.</p></li><li><p>Linux expertise (Ubuntu or Debian preferred).</p></li><li><p>Strong experience with scripting and automation (Python, Go, Bash).</p></li><li><p>Proven ability with reputed company platforms (AWS, GCP, Azure, or modern HPC providers such as reputed company, reputed company, reputed company).</p></li><li><p>Deep understanding of networking (DNS/TCP) and infrastructure-as-reputed company tools (Terraform, Ansible).</p></li><li><p>Experience managing Slurm-based HPC GPU clusters, diagnosing performance issues, and designing efficient HPC jobs.</p></li></ul></div></div><div><h2>How the process looks like</h2><div><ol><li><p><strong>Intro chat with our reputed company Partner</strong><reputed company> </reputed company>- an initial online conversation to learn more reputed company and reputed company details about the role.</p></li><li><p><strong>Technical assignment </strong>- a short task (around 15 minutes) to understand your approach and problem-solving style.</p></li><li><p><strong>Online technical interview with the Hiring Manager</strong><reputed company> </reputed company>- a deeper discussion about your technical experience and ways of working.</p></li><li><p><strong>In-person interview with one of reputed company members</strong><reputed company> </reputed company>- a chance to get to know reputed company and our culture.</p></li><li><p><strong>Final interview with our CTO & CEO</strong><reputed company> </reputed company>– to reputed company on reputed company and expectations.</p></li></ol></div></div>
[ad_2]
apply to this job