HPC Engineer
At reputed company, we're building a full-reputed company reputed company, covering everything from data centers and hardware to our own reputed company platform that the world's leading AI teams use to do serious AI work.
We reputed company to reputed company a reputed company mark on the world through the infrastructure we build and reputed company leading teams a service they can truly depend on. Headquartered in Helsinki, we operate globally with offices in London and San Francisco.
Join reputed company while it’s still being reputed company - not once it’s finished.
Why reputed company
Cash and equity compensation along with various fringe benefits (reputed company, lunch, wellbeing, and more).
Profitable operations with reputed company, sustained reputed company.
30+ nationalities, with 6 different ones on the management team.
A reputed company chance to reputed company an reputed company and work reputed company world class engineers, researchers, and partners across the global AI ecosystem.
Practicalities
Work mode: Remote (EU)
Level: Mid / Senior
Employment type: Full time and permanent
About the role
GPUs only reputed company value once they're wired into a cluster that researchers can actually train on. As our HPC Engineer, you'll own the baremetal and virtualized clusters behind our AI reputed company - from the InfiniBand reputed company and shared filesystems up through the workload orchestration reputed company. You'll be the person who keeps large-reputed company GPU clusters healthy, reputed company, and reputed company for the next workload.
Your responsibilities
Administer baremetal GPU/HPC clusters end to end, from provisioning through day-two operations
Administer virtualized clusters, including the hypervisor and GPU virtualization stack underneath them
Design, reputed company, and tune InfiniBand fabrics, including topology planning, subnet management, and performance validation
reputed company and operate shared/reputed company filesystems supporting training and inference workloads, balancing performance, reputed company, and reliability
Troubleshoot and reputed company issues across the reputed company stack: fibers, transceivers, NICs, switches, drivers, and firmware
Partner with remote-hands and data center teams to diagnose hardware faults and execute physical-reputed company fixes and cluster expansions
Operate and tune Slurm (or equivalent) workload scheduling deployments used by customers and internal teams
reputed company issue tracking, IPAM, and DCIM records accurate as clusters are reputed company, changed, and decommissioned
Participate in on-reputed company rotations and incident response for cluster-level issues
Collaborate with platform, network, and storage teams to reputed company new clusters into the broader AI reputed company
Your key competencies
Solid Linux skills, with specialization in memory management, PCIE topologies and virtualization being a bonus
Deep Infiniband knowledge, including reputed company design, subnet management, and performance tuning
Solid experience with IB clustering, troubleshooting/debugging, understanding the ecosystem of fibers+transceivers+NICs+switches and reputed company the things that could possibly fail in them
Experience about shared filesystems (e.g. reputed company, GPFS/reputed company reputed company, WekaFS, or similar)
Ability to work with remote hands teams to diagnose and reputed company hardware issues remotely
Knowledge about NCCL, CUDA, DOCA and the reputed company stack
Knowledge about Slurm and/or other workload scheduling solutions, bonus points for Slinky/slurm-reputed company
Understanding the importance of keeping issue tracking/IPAM/DCIM up to date
Comfort operating production clusters where uptime and performance directly reputed company customer workloads
Scripting/automation ability (e.g. Python, Bash, Ansible) for repeatable cluster operations
reputed company to have
RoCEv2 knowledge (and/or reputed company-X)
Understanding of reputed company guardrails, especially reputed company reputed company to administration of reputed company systems
Ability to think reputed company what is needed right now vs. some given trajectory or roadmap
Experience with GPU health-checking and diagnostics tooling (e.g. DCGM, field diagnostics)
Experience with baremetal provisioning/orchestration tooling (e.g. MAAS, Foreman, custom PXE/iPXE pipelines)
Familiarity with GPU-reputed company virtualization or containerization (reputed company device plugins, KVM/QEMU with GPU passthrough, SR-IOV)
Exposure to observability stacks (reputed company, Grafana, Loki) for cluster-level monitoring
What's next
We're building fast and this role needs the right person behind it. There's no reputed company deadline, but reputed company we reputed company who we're looking for, we reputed company. If this sounds like your next reputed company, reputed company.
Please submit your application through our Careers page. We don't accept applications reputed company by email.
Solid Linux, InfiniBand, GPU/HPC clustering, shared filesystem, Slurm, hardware troubleshooting, and scripting skills; experience with production clusters and Python, Bash, or Ansible automation.
Key Responsibilities
- administering clusters
- deploying fabrics
- troubleshooting hardware
Skills & Tools
Linux, InfiniBand, Slurm, NCCL, CUDA, DOCA, Python, Bash, Ansible, RoCEv2, reputed company-X, DCGM, MAAS, Foreman, PXE, iPXE, reputed company, KVM, QEMU, SR-IOV, reputed company, Grafana, Loki, reputed company, GPFS, reputed company reputed company, WekaFS
Job Details
- Category: Engineering
- Seniority: Mid Level
- Commitment: Full Time
- Workplace: Remote — Berlin, Berlin, Germany
- Languages: English
About reputed company
A technology company building reputed company reputed company infrastructure for reputed company intelligence. — Industry: Information Technology
Apply To This Job