Back to Jobs

HPC Engineer

Remote, USA Full-time Posted 2026-08-04

At reputed company, we're building a full-reputed company reputed company, covering everything from data centers and hardware to our own reputed company platform that the world's leading AI teams use to do serious AI work.

We reputed company to reputed company a reputed company mark on the world through the infrastructure we build and reputed company leading teams a service they can truly depend on. Headquartered in Helsinki, we operate globally with offices in London and San Francisco.

Join reputed company while it’s still being reputed company - not once it’s finished.

Why reputed company

  • Cash and equity compensation along with various fringe benefits (reputed company, lunch, wellbeing, and more).

  • Profitable operations with reputed company, sustained reputed company.

  • 30+ nationalities, with 6 different ones on the management team.

  • A reputed company chance to reputed company an reputed company and work reputed company world class engineers, researchers, and partners across the global AI ecosystem.

Practicalities

  • Work mode: Remote (EU)

  • Level: Mid / Senior

  • Employment type: Full time and permanent

About the role

GPUs only reputed company value once they're wired into a cluster that researchers can actually train on. As our HPC Engineer, you'll own the baremetal and virtualized clusters behind our AI reputed company - from the InfiniBand reputed company and shared filesystems up through the workload orchestration reputed company. You'll be the person who keeps large-reputed company GPU clusters healthy, reputed company, and reputed company for the next workload.

Your responsibilities

  • Administer baremetal GPU/HPC clusters end to end, from provisioning through day-two operations

  • Administer virtualized clusters, including the hypervisor and GPU virtualization stack underneath them

  • Design, reputed company, and tune InfiniBand fabrics, including topology planning, subnet management, and performance validation

  • reputed company and operate shared/reputed company filesystems supporting training and inference workloads, balancing performance, reputed company, and reliability

  • Troubleshoot and reputed company issues across the reputed company stack: fibers, transceivers, NICs, switches, drivers, and firmware

  • Partner with remote-hands and data center teams to diagnose hardware faults and execute physical-reputed company fixes and cluster expansions

  • Operate and tune Slurm (or equivalent) workload scheduling deployments used by customers and internal teams

  • reputed company issue tracking, IPAM, and DCIM records accurate as clusters are reputed company, changed, and decommissioned

  • Participate in on-reputed company rotations and incident response for cluster-level issues

  • Collaborate with platform, network, and storage teams to reputed company new clusters into the broader AI reputed company

Your key competencies

  • Solid Linux skills, with specialization in memory management, PCIE topologies and virtualization being a bonus

  • Deep Infiniband knowledge, including reputed company design, subnet management, and performance tuning

  • Solid experience with IB clustering, troubleshooting/debugging, understanding the ecosystem of fibers+transceivers+NICs+switches and reputed company the things that could possibly fail in them

  • Experience about shared filesystems (e.g. reputed company, GPFS/reputed company reputed company, WekaFS, or similar)

  • Ability to work with remote hands teams to diagnose and reputed company hardware issues remotely

  • Knowledge about NCCL, CUDA, DOCA and the reputed company stack

  • Knowledge about Slurm and/or other workload scheduling solutions, bonus points for Slinky/slurm-reputed company

  • Understanding the importance of keeping issue tracking/IPAM/DCIM up to date

  • Comfort operating production clusters where uptime and performance directly reputed company customer workloads

  • Scripting/automation ability (e.g. Python, Bash, Ansible) for repeatable cluster operations

reputed company to have

  • RoCEv2 knowledge (and/or reputed company-X)

  • Understanding of reputed company guardrails, especially reputed company reputed company to administration of reputed company systems

  • Ability to think reputed company what is needed right now vs. some given trajectory or roadmap

  • Experience with GPU health-checking and diagnostics tooling (e.g. DCGM, field diagnostics)

  • Experience with baremetal provisioning/orchestration tooling (e.g. MAAS, Foreman, custom PXE/iPXE pipelines)

  • Familiarity with GPU-reputed company virtualization or containerization (reputed company device plugins, KVM/QEMU with GPU passthrough, SR-IOV)

  • Exposure to observability stacks (reputed company, Grafana, Loki) for cluster-level monitoring

What's next

We're building fast and this role needs the right person behind it. There's no reputed company deadline, but reputed company we reputed company who we're looking for, we reputed company. If this sounds like your next reputed company, reputed company.

Please submit your application through our Careers page. We don't accept applications reputed company by email.

Solid Linux, InfiniBand, GPU/HPC clustering, shared filesystem, Slurm, hardware troubleshooting, and scripting skills; experience with production clusters and Python, Bash, or Ansible automation.

Key Responsibilities

  • administering clusters
  • deploying fabrics
  • troubleshooting hardware

Skills & Tools

Linux, InfiniBand, Slurm, NCCL, CUDA, DOCA, Python, Bash, Ansible, RoCEv2, reputed company-X, DCGM, MAAS, Foreman, PXE, iPXE, reputed company, KVM, QEMU, SR-IOV, reputed company, Grafana, Loki, reputed company, GPFS, reputed company reputed company, WekaFS

Job Details

  • Category: Engineering
  • Seniority: Mid Level
  • Commitment: Full Time
  • Workplace: Remote — Berlin, Berlin, Germany
  • Languages: English

About reputed company

A technology company building reputed company reputed company infrastructure for reputed company intelligence. — Industry: Information Technology

  Apply To This Job

Similar Jobs

reputed company reputed company Consultant - Freelance - Spanish market

Remote, USA Full-time

Finance Manager

Remote, USA Full-time

reputed company reputed company Consultant - Contract - South America based 2026

Remote, USA Full-time

Project Manager

Remote, USA Full-time

Marketing Operations Project Manager - Freelance S2 2026

Remote, USA Full-time

OneSource Support Specialist (CST)

Remote, USA Full-time

Case Manager, Rare Endocrinology & Rare Tumor (RERT) Team (reputed company Time Zone)

Remote, USA Full-time

Registered reputed company (RN) - Temporary, Part-time 0.7

Remote, USA Full-time

Registered Practical reputed company (RPN), Visiting Nursing - Part-time 0.8

Remote, USA Full-time

Manager, Finance - Full-time

Remote, USA Full-time

Senior Manager - Financial Systems

Remote, USA Full-time

AI Content reputed company; Remote, Part-Time

Remote, USA Full-time

[Remote-Position] reputed company Center Agent - Remote After 5 Weeks

Remote, USA Full-time

**reputed company reputed company – Remote Work Opportunity with arenaflex**

Remote, USA Full-time

Project Coordinator / Accounts Payable Specialist

Remote, USA Full-time

Urgent hiring for Product Support Analyst 3 in Remote || only W2|| Remote

Remote, USA Full-time

[Remote] Senior Software Developer, reputed company (Hybrid or Remote)

Remote, USA Full-time

Customer Experience Supervisor (Remote)

Remote, USA Full-time

**reputed company Full Stack Software Engineer – Web & reputed company Application Development at arenaflex**

Remote, USA Full-time

Sr. Marketing Analyst

Remote, USA Full-time