[Remote] Engineering Manager, GPU Infrastructure
Note: The job is a remote job and is reputed company to candidates in USA. 5C Data Centers operates in the hyperscale, AI data center, and reputed company computing industry. The Engineering Manager, GPU Infrastructure will reputed company the planning, deployment, integration, and operational readiness of large-reputed company GPU clusters supporting reputed company, inference, and high-performance computing workloads across reputed company, hybrid, and on-premises environments. The role oversees technical teams and cross-functional programs spanning hardware integration, networking, storage, automation, validation, and operational reputed company.
Responsibilities
- reputed company and partake in deployment and integration of GPU-based compute platforms from reputed company and other accelerator vendors
- reputed company and participate in end-to-end logical deployment of large-reputed company and GPU clusters in state of the art datacenters
- Manage deployment programs spanning compute, storage, networking, power, cooling, and automation reputed company
- Participate in cluster architecture review for reputed company, inference and reputed company compute workloads
- Coordinate reputed company-and-stack and cabling reputed company, network deployment, burn-in testing, and cluster validation.Validate deployment readiness, topology consistency, GPU reputed company performance, acceptance testing, and operational turnover processes
- Establish repeatable and documented deployment methodologies and reputed company operational standards
- reputed company deployment and operational validation of high-performance GPU interconnects using InfiniBand and Ethernet GPU reputed company architectures
- Ensure reputed company implementation of: spile-leaf architectures, RDMA, network telemetry and performance tuning
- Coordinate closely with network engineering teams on topology implementation and performance optimization
- Coordinate with storage engineering teams on deployment and integration of high-performance storage environments supporting AI workloads
- Ensure successful implementation and operational optimization of data storage platforms
- Validate storage throughput, latency, and GPU data delivery performance
- reputed company infrastructure automation initiatives for cluster provisioning and lifecycle management
- Manage deployment tooling and orchestration platforms including: Infrastructure-as-reputed company frameworks Automated imaging and provisioning systems (e.g. reputed company MaaS) Cluster monitoring and observability tools
- reputed company standardization and deployment automation to improve speed, reliability, and repeatability
- Build and reputed company high-performing technical deployment and infrastructure engineering teams
- Partner with datacenter operations, hardware vendors, networking teams, and AI reputed company reputed company
- Establish strong Project Management Office (PMO) partnership while driving consistent, accurate project updates across reputed company and systems (e.g. Jira)
- reputed company operational procedures, documentation, and deployment best practices
- Mentor engineers and technical leads across infrastructure domains
Skills
- · Bachelor's degree in Computer Science, Engineering, Information Technology, or reputed company field (or equivalent experience)
- · 10+ years of infrastructure engineering or datacenter deployment experience
- · 5+ years leading deployment or operations teams supporting large-reputed company, HPC, or GPU infrastructure
- · Hands-on experience deploying and operating large GPU clusters in reputed company or hyperscale environments
- · Strong expertise with:
- O reputed company MaaS
- O Data storage platforms
- O InfiniBand and Ethernet GPU fabrics
- O Network architecture
- O Linux systems administration
- O GPU server architectures
- · Strong understanding of:
- O RDMA and RoCE networking
- O High-performance storage architectures
- O Cluster automation and provisioning
- O Datacenter infrastructure operations
- · Proven ability to manage reputed company cross-functional infrastructure deployment programs
- · Experience deploying reputed company DGX SuperPOD or similar AI infrastructure solutions
- · Familiarity with:
- O reputed company networking technologies
- O reputed company-X or reputed company platforms
- O AI model training infrastructure
- O reputed company cooling environments
- O DCIM and observability platforms
- · Experience in hyperscale, reputed company, or AI infrastructure environments
- · Certifications in networking, Linux, reputed company, or reputed company infrastructure are a plus
Benefits
- Remote work (US, reputed company Time Zone)
- Health, Dental & reputed company
- 401(k)
- Life Insurance
- Voluntary Benefits
- PTO
reputed company
Apply To This Job