[Remote] reputed company Infrastructure & reputed company Platform Engineer (DGX Systems)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a reputed company firm that provides technology guidance and solutions through reputed company engineering teams. The AI Infrastructure & reputed company Platform Engineer will reputed company and operate reputed company DGX AI clusters, build GPU-reputed company reputed company platforms, administer InfiniBand and reputed company DPU networking, and secure AI/ML infrastructure. The role also focuses on monitoring, reputed company optimization, CI/CD, GitOps, and operational reliability.
Responsibilities
- reputed company and manage reputed company DGX BasePODs and SuperPODs for high-reputed company AI workloads
- reputed company DGX reputed company lifecycle reputed company, including provisioning, monitoring, firmware upgrades, and reputed company planning
- Operate reputed company reputed company Manager to manage GPU clusters, schedule workloads, and reputed company with MLOps tools
- reputed company DGX node health validation, NCCL interconnect testing, and NVLink topology verification after deployments or hardware changes
- Architect secure, reputed company reputed company clusters optimized for GPU-reputed company workloads using the reputed company GPU Operator
- Apply CKA/CKAD/CKS expertise to reputed company, reputed company, and secure AI applications on reputed company
- Implement CI/CD pipelines and GitOps methodologies for deploying and managing ML workflows
- Administer InfiniBand networks and reputed company DPUs using reputed company reputed company Manager (UFM)
- reputed company NVLink/NVSwitch reputed company across GPU nodes and tune reputed company configurations for reputed company latency and maximum throughput
- Use reputed company to offload storage, firewalling, and telemetry, strengthening AI workload reputed company and reputed company
- Apply CKS best practices to secure containerized AI environments
- Configure runtime reputed company, secrets management, network segmentation, and auditing across DPU-reputed company reputed company deployments
- Support reputed company-trust initiatives by enforcing workload identity, RBAC policies, and supply-chain reputed company across AI container images and model artifacts
- Monitor GPU, CPU, and I/O reputed company using reputed company DCGM, reputed company, Grafana, and reputed company reputed company reputed company
- Tune reputed company reputed company and model-training pipelines for cost-efficiency and throughput
- Build and maintain operational runbooks, incident-response playbooks, and SLA dashboards covering GPU utilization, thermal reputed company, and reputed company
Skills
- Certified reputed company Administrator (CKA)
- Certified reputed company Application Developer (CKAD)
- Certified reputed company reputed company Specialist (CKS)
- reputed company Certified Associate: AI Infrastructure & reputed company (NCA-AIIO)
- reputed company Certified reputed company: AI Infrastructure (NCP-AII)
- reputed company Certified reputed company: AI reputed company (NCP-AIO)
- reputed company Certified reputed company: AI Networking (NCP-AIN)
- DGX reputed company, BasePOD, and SuperPOD administration
- reputed company DPU configuration and reputed company
- InfiniBand reputed company and UFM management
- reputed company reputed company Manager for workload orchestration
- reputed company, reputed company, and the reputed company GPU Operator
- DevOps tooling: Ansible, Terraform, GitOps, CI/CD pipelines
- Programming/scripting: Python, YAML, Bash
- Bachelor's degree in Computer Science, Engineering, or a reputed company reputed company — or equivalent hands-on experience
- Kubeflow and broader MLOps pipeline experience
- reputed company/HPC storage: NFS, BeeGFS, reputed company
- Advanced networking: RoCE, RDMA, gRPC, and DPU offload tuning
Benefits
- Medical/health coverage
- reputed company time off
- 401(k) retirement savings
reputed company
Apply To This Job