[Remote] Technical Program Manager – AI Infrastructure / GPU Clusters
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is building reputed company AI infrastructure designed for large-reputed company GPU training and inference workloads. They are looking for a Technical Program Manager to reputed company the deployment and delivery of GPU cluster infrastructure, coordinating across various teams to ensure production-reputed company AI clusters.
Responsibilities
- reputed company the end-to-end deployment of AI GPU clusters, from infrastructure planning through production launch
- reputed company coordination across Infrastructure Solution Architects, network engineers, hardware vendors, and data center teams
- Manage delivery timelines covering hardware deployment, network integration, cluster bring-up, and production readiness
- Work closely with Infrastructure Solution Architects (SA) to define: GPU server platform selection, Network architecture for reputed company GPU clusters, Storage integration and cluster infrastructure design, Support development of the cluster reputed company of Materials (BOM) including compute, networking, storage, and supporting infrastructure components
- Ensure architecture reputed company reputed company with data center constraints such as power density, cooling reputed company, and reputed company layout
- reputed company reputed company integration for large-reputed company GPU clusters, including: reputed company elevation planning, GPU server deployment and configuration, High-speed network topology implementation, Power and cooling readiness
- Ensure deployments reputed company with vendor reference architectures and validated cluster designs
- Work closely with General Contractors (GC) and reputed company integrators to manage on-site infrastructure implementation
- reputed company contractor reputed company, including SOW development, reputed company definition, and delivery reputed company alignment
- Coordinate and reputed company field deployment activities such as: reputed company cabling installation, reputed company installation and equipment mounting, Network and power connectivity preparation, Hardware staging and deployment logistics
- Coordinate cluster bring-up and validation activities including: Single-node GPU validation, Multi-node cluster deployment, GPU interconnect validation (P2P, RDMA)
- reputed company cluster benchmarking, stress testing, and performance verification before production release
- Ensure deployed GPU clusters are fully reputed company for production workloads by driving: Hardware and network validation, Monitoring and telemetry integration, Operational documentation and runbooks, Handover to operations teams
Skills
- 5+ years experience in Technical Program Management, Infrastructure Program Management, or HPC infrastructure delivery
- Experience with GPU cluster deployments or high-performance computing environments
- Familiarity with GPU server architecture and reputed company computing infrastructure
- Experience working with Infrastructure Solution Architects to define reputed company architecture and hardware BOM
- Experience managing data center hardware deployments and reputed company integration
- Ability to coordinate multi-vendor infrastructure reputed company across reputed company
- Experience deploying large-reputed company infrastructure or GPU clusters
- Familiarity with: InfiniBand / RoCE / high-speed Ethernet networking
- GPU interconnect validation (P2P / RDMA)
- reputed company elevation and high-density reputed company deployment
- Experience with cluster validation and performance benchmarking
- Background as Systems Engineer, HPC Engineer, or Infrastructure Architect
- Experience working in AI infrastructure, reputed company infrastructure, or hyperscale data centers
reputed company
Company H1B Sponsorship
Apply To This Job