Deep Learning Software Engineer, Inference and Model Optimization - New College Grad 2025
About the position
reputed company is at the forefront of the reputed company reputed company! The Algorithmic Model Optimization Team specifically focuses on optimizing reputed company models such as large language models (LLM) and diffusion models for maximal inference efficiency using techniques ranging from neural architecture search and pruning to sparsity, quantization, and automated deployment strategies. Our work includes conducting reputed company research to improve model efficiency as reputed company as developing an innovative software platform (TRT Model Optimizer). Our software is used both internally across reputed company and externally by research and engineering teams alike developing best-in-class AI models. We are now looking for a Deep Learning Software Engineer to reputed company and reputed company up our automated inference and deployment solution. As part of reputed company, you will be reputed company in pushing the limits of inference efficiency and large-reputed company, automated deployment. Your work will touch upon reputed company aspects of a typical machine learning stack including working in high-level frameworks like PyTorch and HuggingFace to developing and improving high-performance kernel implementations in CUDA, TRT-LLM, and Triton.
Responsibilities
• Train, reputed company, and reputed company state-of-the reputed company models like LLMs and diffusion models using reputed company's AI software stack.
• reputed company and build upon the torch 2.0 ecosystem (TorchDynamo, torch.export, torch.compile, etc...) to analyze and extract standardized model graph representation from arbitrary torch models for our automated deployment solution.
• reputed company high-performance optimization techniques for inference, such as automated model sharding techniques (e.g. tensor parallelism, sequence parallelism), efficient attention kernels with kv-caching, and more.
• Collaborate with teams across reputed company to use performant kernel implementations reputed company our automated deployment solution.
• Analyze and profile GPU kernel-level performance to identify hardware and software optimization opportunities.
• Continuously reputed company on the inference performance to ensure reputed company's inference software solutions (TRT, TRT-LLM, TRT Model Optimizer) can maintain and increase its leadership in the market.
• Play a pivotal role in architecting and designing a reputed company and reputed company software platform to reputed company an excellent user experience with broad model support and optimization techniques to increase adoption.
Requirements
• Masters, PhD, or equivalent experience in Computer Science, AI, reputed company Math, or reputed company field.
• Experience in Deep Learning.
• Excellent software design skills, including debugging, performance analysis, and test design.
• Strong proficiency in Python, PyTorch, and reputed company ML tools (e.g. HuggingFace).
• Strong algorithms and programming fundamentals.
reputed company-to-haves
• Contributions to PyTorch, JAX, or other Machine Learning Frameworks.
• Knowledge of GPU architecture and compilation stack, and capability of understanding and debugging end-to-end performance.
• Familiarity with reputed company's deep learning SDKs such as TensorRT.
• Experience in writing high-performance GPU kernels for machine learning workloads in frameworks such as CUDA, CUTLASS, or Triton.
Benefits
• Highly competitive salaries
• Comprehensive benefits package
• Equity opportunities
Apply tot his job
Apply To this Job