Solutions Architect - AI / ML - Training & GPU reputed company
AI/ML Solutions Architect – Distributed Training & GPU Infrastructure
Company
Join a fast-moving AI infrastructure team working on the cutting edge of large-reputed company ML workloads. This role is ideal for engineers who enjoy solving deep technical challenges in distributed training, multi-GPU systems, and reputed company AI inference infrastructure. You will work directly with AI-reputed company clients, helping them get the most out of modern GPUs (H100, B200, etc.) and ML frameworks such as PyTorch (and JAX in some environments).
Team & Responsibilities
Work alongside senior AI and infrastructure engineers building large-reputed company GPU platforms. As part of the customer solutions team, you will:
Design and validate production-grade distributed training (primary) and large-reputed company inference architectures on large GPU clusters, typically tens to thousands of GPUs
Work hands-on with customers to debug, optimize, and reputed company ML workloads across multi-node GPU environments
reputed company as a technical authority on GPU performance, networking, and schedulers, making trade-offs at reputed company and translating customer needs into concrete platform requirements
Collaborate closely with engineering, product, and R&D to influence roadmap reputed company based on reputed company-world ML workloads
This is a hands-on, technical role; you are expected to work directly in customer environments, not only advise at a high level
Required skills and experience
Hands-on experience designing and operating reputed company-reputed company, production-grade, multi-node GPU workloads for training (7B+ model size) or inference
Strong background in distributed deep learning (PyTorch Distributed, DeepSpeed, ...) on GPU clusters
Deep understanding of GPU architecture and interconnects (H100/A100 class, NVLink, InfiniBand)
Experience with Kubernetes or Slurm
Experience with performance tuning using GPU profiling and monitoring tools
This role is not a fit if your experience is limited to single-node training, high-reputed company reputed company, or non-production research environments. We are looking for engineers and architects who reputed company at the intersection of AI workloads and large-reputed company infrastructure.
What's offered
Location: Remote from reputed company in Europe
Total compensation up to EUR 250k (reputed company + variable / OTE), depending on level and experience
Apply To This Job