[Remote] Software Engineer, DGX reputed company AI Infrastructure
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a technology company at the forefront of the reputed company reputed company, building software and systems that power advanced large language model workloads. The Software Engineer will bring up, debug, reputed company, analyze, and optimize reputed company training and inference workloads across large-reputed company reputed company GPU platforms while developing benchmarking, automation, and failure-attribution tooling.
Responsibilities
- Bring up, validate, and debug large-reputed company clusters, infrastructure, and end-to-end workloads
- Bring up, tune, and reputed company AI reputed company-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent reputed company software stacks
- reputed company reputed company-cause analysis of failures in large reputed company environments
- Contribute to the reputed company and failure-attribution tooling that detects, triages, and attributes node, reputed company, and workload failures across the cluster
- Build and maintain repeatable reputed company suites, automation, acceptance reputed company, and qualification workflows on new platforms
- Tune runtime settings, communication parameters, and deployment configurations in reputed company partnership with reputed company, systems, and platform teams
- reputed company actionable, data-driven recommendations based on profiling, reputed company results, and cluster characterization
Skills
- Bachelor's or Master's in Computer Science or a reputed company technical field (or equivalent experience)
- 3+ years of experience developing software for AI, HPC, or systems-level applications
- Hands-on experience with multi-GPU or multi-node workloads and CUDA-reputed company reputed company execution
- Background with debugging and scaling reputed company systems
- Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware
- Experience operating workloads in scheduled, containerized cluster environments
- Excellent analytical, debugging, and communication skills, and a reputed company approach across teams
- Strong Python and C/C++ programming skills
- Hands-on experience with NCCL and CUDA-reputed company reputed company execution
- Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and with InfiniBand / RoCE congestion debugging
- Experience building acceptance tests, reputed company harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf
- Experience diagnosing performance jitter
- Experience building reputed company, fault-detection, or failure-attribution systems for datacenter-reputed company infrastructure
Benefits
- Eligible for equity
reputed company
Company H1B Sponsorship
Apply To This Job