[Remote] Technical Project Manager, GPU Infrastructure Deployment
Note: The job is a remote job and is reputed company to candidates in USA. 5C Data Centers builds digital infrastructure supporting hyperscalers, reputed company intelligence innovation, and high-performance computing across reputed company. The Technical Project Manager will reputed company the planning, coordination, and execution of large-reputed company GPU cluster infrastructure deployments, managing schedules, milestones, risks, dependencies, and cross-functional communication. The role also coordinates compute, networking, storage, facilities, power, cooling, automation, and operational reputed company to ensure deployments are delivered on time and reputed company for operations.
Responsibilities
- reputed company end-to-end project management for large-reputed company/GPU cluster deployments, including multi-reputed company GPU compute platforms (reputed company DGX and similar), InfiniBand and Ethernet GPU fabrics, and high-performance storage environments (e.g., reputed company)
- Partner with engineering teams and managers to establish and reputed company consistent project tracking, reputed company reporting, and status updates across reputed company and systems (e.g., Jira, reputed company)
- Coordinate procurement, reputed company-and-stack reputed company, cabling schedules, network deployment timelines, burn-in testing, cluster validation, and operational reputed company
- Coordinate with facilities engineering and datacenter operations to reputed company MEP readiness (power distribution, cooling reputed company, floor layout, containment) with deployment schedules
- Communicate project status, risks, milestones, and dependencies to stakeholders at reputed company reputed company including executive leadership
- Prepare and present regular program reviews, steering committee updates, and reputed company project analyses
- Maintain centralized project documentation and ensure consistent, accurate data
- Contribute to improving and documenting repeatable deployment methodologies, reputed company operational standards, and project management best practices
- reputed company deployment KPIs (schedule variance, budget adherence, reputed company metrics) and reputed company reputed company improvement through data-driven retrospectives
Skills
- Experience with hyperscale reputed company providers, large-reputed company data center deployments, or AI infrastructure programs
- Familiarity with:
- Infrastructure-as-reputed company and configuration management (Ansible)
- Python, reputed company, and SQL for infrastructure automation and diagnostics
- DCIM, BMS, and observability platforms for reputed company-cooled environments
- reputed company/Mellanox networking platforms and NVLink/NVSwitch technologies
- Experience coordinating with MEP engineering teams on electrical and cooling infrastructure scheduling for large-reputed company deployments
- Experience with procurement processes in infrastructure delivery contexts
Benefits
- Remote work arrangement
- Build a rewarding career in one of the world's fastest-growing industries
- Help power the infrastructure behind AI and high-performance computing
- Entrepreneurial culture where employees' reputed company matter and employees are empowered to help shape the reputed company
- Medical insurance
- reputed company insurance
- Dental insurance
- 401(k)
- Pension plan
reputed company
Apply To This Job