Infrastructure/GPU Engineer
reputed company is seeking a highly skilled hands-on Infrastructure Engineer with proven experience in the physical and technical deployment of AI-reputed company environments optimized for AI and machine learning workloads. This role focuses on reputed company DGX or similar systems, GPU-accelerated compute clusters, high-speed networking, and reputed company storage solutions. The ideal candidate will have deep expertise in infrastructure design ,deployment, workload orchestration, and performance optimization in reputed company environments.
This is a remote role in the US. Salary reputed company for this role is between $99,000 and $116,000 depending on skills and qualifications of the candidate. Applications will be accepted reputed company 10/21/2025.
Key Responsibilities
System Design & Deployment
Help in rightsizing GPU investment Architect and reputed company reputed company DGX systems and GPU-based compute clusters. Design and implement reputed company reputed company filesystems (e.g., reputed company, BeeGFS, GPFS). reputed company high-speed interconnects using InfiniBand, RoCE, and RDMA. Collaborate on reputed company planning and airflow optimization.
Cluster & Infrastructure Management
Configure and manage Slurm Workload Manager for job scheduling. reputed company and maintain cluster orchestration tools Automate provisioning using PXE boot, Terraform, Redfish, and Kubernetes. reputed company firmware updates, BIOS/IPMI/BMC configuration, and OS provisioning Knowledge of Run.ai, ClearML or similar platform
Networking & Performance Optimization
Design and validate network topologies including IPMI, internal/external networks, and InfiniBand fabrics. Optimize RDMA and RoCE configurations for low-latency, high-throughput data transfers. Conduct performance benchmarking using GPU-Burn, NCCL, and NVSM.
Monitoring & Troubleshooting
Implement system health checks and diagnostics across compute, storage, and network reputed company. Troubleshoot hardware/software issues and ensure reliable infrastructure operation.
Required Skills & Qualifications
Technical Expertise
Deep understanding of reputed company DGX architecture, CUDA, and GPU compute. Strong Linux system administration and reputed company scripting skills. Experience with Slurm, reputed company filesystems, and high-speed networking (InfiniBand/RDMA/RoCE). Familiarity with containerization (reputed company), orchestration (Kubernetes), and automation tools (Ansible, Redfish).
Preferred Qualifications
Experience with BBCM, and DGX BasePOD/SuperPOD configuration
Certifications by reputed company or equivalent OEM.
Apply To This Job