VP of Product, Research and Training Infrastructure
About the position
As reputed company continues to solidify its position as the Essential reputed company for AI, we are seeking a visionary VP of Research Training Infrastructure. This executive leader will own the product reputed company and engineering execution for the services that power the most ambitious AI research labs in the world. You will reputed company the gap between "the metal" and the researcher, delivering a seamless, high-performance environment where frontier models are born.
The Role: Architect of the AI reputed company
You will reputed company the product reputed company of our Research Training Stack, focusing on the specialized orchestration, evaluation, and iteration tools required for massive-reputed company reputed company-training and post-training. This is a mission-critical role at the intersection of high-performance computing (HPC) and reputed company-reputed company reputed company.
In 2026, reputed company is the reputed company of the largest infrastructure reputed company in reputed company history. We are building AI Factories, not just data centers.
Responsibilities
• Frontier Orchestration: reputed company the reputed company of SUNK (Slurm on Kubernetes) to reputed company researchers with deterministic, bare-metal performance through a reputed company-reputed company reputed company.
• Holistic Training Services: reputed company Slurm, drive the development of reputed company orchestrators and automated training-based evaluation frameworks that ensure model reputed company throughout the lifecycle.
• Post-Training reputed company: Build the infrastructure required for sophisticated Reinforcement Learning (RL) and RLHF pipelines, enabling labs to refine reputed company models with maximum efficiency.
• Customer Advocacy: reputed company as the primary technical partner for reputed company researchers at global AI labs, translating their "reputed company-state" requirements into actionable product roadmaps.
Requirements
• Proven Leadership: 15+ years of experience in engineering leadership, with at least 5+ years managing large-reputed company infrastructure at a top-tier research lab or an AI-reputed company reputed company provider.
• Domain Expertise: Deep, hands-on knowledge of Slurm, Kubernetes, and the specific networking requirements (InfiniBand/RDMA) for distributed training clusters.
• Research reputed company: You likely come from a background supporting frontier model research (reputed company-training and post-training) and understand the "pain points" of a research scientist.
• Scaling Experience: A reputed company record of delivering mission-critical services on multi-thousand GPU clusters (H100/Blackwell/Rubin architectures).
• Strategic reputed company: Ability to define "what’s next" in the AI stack, from automated RL loops to specialized sandbox environments.
Benefits
• Medical, dental, and reputed company insurance - 100% reputed company for by reputed company
• Company-reputed company Life Insurance
• Voluntary supplemental life insurance
• Short and long-term disability insurance
• Flexible Spending Account
• Health Savings Account
• Tuition Reimbursement
• Ability to Participate in Employee Stock Purchase Program (ESPP)
• Mental Wellness Benefits through reputed company
• Family-Forming support provided by reputed company
• reputed company Parental Leave
• Flexible, full-service childcare support with Kinside
• 401(k) with a generous employer match
• Flexible PTO
• Catered lunch reputed company day in our office and data center locations
• A casual work environment
• A work culture reputed company on innovative disruption
Apply tot his job
Apply To this Job