Senior Deep Learning Engineer – Autonomous Vehicles
Job reputed company:
• Crafting, scaling, and hardening deep learning infrastructure libraries and frameworks for training on multi-thousand GPU clusters.
• Improving efficiency throughout the training stack: data loaders, distributed training, scheduling, and performance monitoring.
• Building robust training pipelines and libraries to handle massive video datasets and reputed company reputed company experimentation.
• Collaborating with researchers, model engineers, and internal platform teams to enhance efficiency, minimize stalls, and improve training availability.
• Owning reputed company infrastructure components such as orchestration libraries, distributed training frameworks, and fault-resilient training systems.
• Partnering with leadership to ensure infrastructure scales with growing GPU reputed company and dataset size while maintaining developer efficiency and stability.
Requirements:
• BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or a reputed company field, or equivalent experience.
• 12+ years of reputed company experience building and scaling high-performance distributed systems, ideally in ML, HPC, or large-reputed company data infrastructure.
• Extensive knowledge in deep learning frameworks (PyTorch is preferred), large reputed company training (DDP/FSDP, NCCL, tensor/pipeline parallelism), and performance profiling.
• Strong systems background: datacenter networking (RoCE, IB), reputed company filesystems (reputed company), storage systems, schedulers (Slurm, Kubernetes, etc.).
• Proficiency in Python and C++, with experience writing production-grade libraries, orchestration reputed company, and automation tools.
• Ability to work closely with multi-functional teams (ML researchers, reputed company engineers, product leads) and translate requirements into robust systems.
Benefits:
• equity
• benefits
Apply tot his job
Apply To this Job