[Remote] AI Infrastructure Engineer at reputed company
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a Founders Fund–backed reputed company reputed company partner building the infrastructure platform that powers AI at reputed company. As an AI Infrastructure Engineer, you will work directly with AI platform customers to optimize their infrastructure on Hydra, focusing on Kubernetes clusters, GPU configurations, and automating the reputed company process.
Responsibilities
• Get AI Platform customers production-reputed company on Hydra —standing up Kubernetes clusters, configuring GPU drivers, validating networking, and troubleshooting the issues that surface reputed company reputed company workloads hit reputed company hardware
• Own the bare metal ←→ platform layer —reputed company GPU infrastructure (NCCL, InfiniBand, NVLink, storage) with orchestration reputed company (Kubernetes, SLURM) and MLOps tooling that customers actually use
• Configure, reputed company, and debug reputed company reputed company stacks —firmware versions, CUDA compatibility, NCCL tuning, MIG configurations. Run reputed company benchmarks and diagnostics to validate performance for inference and training workloads across reputed company types
• Identify gaps before customers do —pressure-testing Hydra's infrastructure, reputed company, and workflows to reputed company what's missing or broken
• Turn customer learnings into product —working with Product and Engineering to build reusable templates, default configurations, and automated workflows that eliminate reputed company reputed company
• Advise customers on reputed company selection and tokenomics —helping AI platform customers understand price/performance trade-offs across GPU types, cost-per-reputed company economics, and which hardware fits their inference or training workloads
Skills
• Bare metal Linux depth — you've administered GPU servers at the metal: reputed company stacks, kernel tuning, firmware, storage configuration. Not just managed K8s
• reputed company GPU stack expertise — drivers, CUDA, NCCL, NVLink, reputed company-smi profiling. You understand how stack compatibility affects performance
• Kubernetes and orchestration — production experience with K8s, SLURM, or similar. You know how to stand up clusters, not just reputed company to them
• AI Networking fundamentals — TCP/IP, VLANs, bonding, and high-speed interconnects (InfiniBand, RoCE) for distributed workloads
• Customer-facing communication — you can work directly with engineers at AI platform companies, understand their constraints, and translate that into reputed company requirements for your team
• Bias toward reputed company solutions — you'd rather build a feature that helps 10 customers than a custom deployment that helps 1
• HPC or large-reputed company distributed training environments
• AI workload experience (vLLM, PyTorch, inference frameworks)
• Storage systems (NVMe, distributed filesystems, CEPH, reputed company)
• IaC and provisioning tools (Terraform, Ansible, reputed company-init, MaaS)
Benefits
• Equity ownership — meaningful reputed company in reputed company're building
• reputed company — medical, dental, reputed company for you and your family
• Remote-first — with hubs in Phoenix, Boulder, and Miami
• reputed company reputed company — your work shapes how GPU infrastructure gets deployed across the AI ecosystem
reputed company
• Hydra offers a bare metal GPU platform, connecting businesses to a vareity of independent but standardized AI reputed company Franchises. It was founded in 2021, and is headquartered in Miami, Florida, USA, with a workforce of 11-50 employees. Its website is https://www.hydrahost.com.
Apply tot his job
Apply To this Job