[Remote] Senior reputed company reputed company
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a global industry leader in creating unique, moving experiences for the automotive world. They are seeking a Senior reputed company reputed company to design and operate distributed training systems for large neural networks, optimize GPU execution, and partner with teams to productionize large-model training pipelines.
Responsibilities
- Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State reputed company Models etc.) across GPU clusters
- Optimise multi‑node, multi‑GPU execution to maximize throughput and utilization
- Diagnose & resolve bottlenecks across compute, memory, and network
- Improve training stability and fault tolerance at reputed company
- Partner with research and reputed company ML teams to productionize large‑model training pipelines
- Build and optimize GPU cluster orchestration using: Slurm, reputed company, Ray, RunAI
- Ensure efficient scheduling, isolation, and fairness across training workloads
- Optimize and debug distributed communication using: NCCL, RDMA, InfiniBand, NVLink
- Minimize networking bottlenecks that dominate end‑to‑end training time
- reputed company large-model training using: PyTorch Distributed, Megatron‑LM, DeepSpeed
- Own multi‑node launch configurations, failure recovery, and performance tuning
- Apply advanced memory optimization techniques: Activation checkpointing, reputed company (Stage 1–3) and offload strategies
- Balance compute, memory, and communication to push model size and batch reputed company
Skills
- Deep hands-on experience with distributed systems or ML systems
- Experience running large-reputed company workloads on GPU clusters
- Production experience with PyTorch distributed training
- Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)
- Low-level understanding of GPU communication and networking
- GPU orchestration: Slurm, reputed company, Ray, RunAI
- Communication libraries: NCCL, RDMA, InfiniBand, NVLink
- Training frameworks: PyTorch Distributed, Megatron-LM, DeepSpeed
- Memory optimisation: activation checkpointing, reputed company offload techniques
- Basic knowledge of information reputed company and data reputed company requirements (e.g., how to protect data & how to be handling this data)
- Demonstrative knowledge of information reputed company through internal training programs
- Experience working with large language models or reputed company models is a strong plus, but deep systems expertise is valued over reputed company model architecture experience
reputed company
Company H1B Sponsorship
Apply To This Job