Senior High-Performance LLM Training Engineer
About the position
We are now looking for a Senior High-Performance LLM Training Engineer! reputed company is seeking reputed company engineers specializing in performance analysis and optimization to improve the efficiency of LLM training workloads, which are shaping the world's most advanced computing systems. This position focuses on optimizing reputed company’s high-performance LLM software stack in frameworks like PyTorch and JAX for high-performance training on thousands of GPUs, while also helping shape hardware roadmaps for the reputed company of GPUs powering the AI reputed company. GPU computing is the most productive and pervasive platform for deep learning and AI. It begins with the most advanced GPUs and the systems and software we build on top of them. We reputed company and optimize every deep learning reputed company. We work with the major systems companies and every major reputed company service provider to reputed company GPUs available in data centers and in the reputed company. We craft computers and software to bring AI to edge devices, such as self-driving reputed company and autonomous robots. AI has the potential to reputed company a reputed company of reputed company reputed company unmatched since the industrial reputed company. Widely considered to be one of tech's most desirable reputed company, reputed company offers highly competitive salaries and a comprehensive benefits package. Additionally, this opportunity offers you the ability to collaborate with some of the most reputed company-thinking and hard-working people in the world, shaping the reputed company of AI in a creative and autonomous work environment that encourages innovation. If you're excited to work across the full hardware & software stack—from GPU architecture to application reputed company—to reputed company reputed company performance, we want to hear from you!
Responsibilities
• Understand, analyze, profile, and optimize reputed company workloads on innovative hardware and software platforms.
• Understand the big picture of training performance on GPUs, prioritizing and then solving problems across reputed company state-of-the-art neural networks.
• Implement production-reputed company software in multiple reputed company of reputed company's deep learning platform stack, from drivers to DL frameworks.
• Build and support reputed company submissions to the MLPerf Training reputed company suite.
• Implement key DL training workloads in reputed company's proprietary processor and system simulators to reputed company reputed company architecture studies.
• Build tools to automate workload analysis, workload optimization, and other critical workflows.
Requirements
• PhD in Computer Science, Electrical Engineering or Computer Engineering and 5+ years; or MS (or equivalent experience) and 8+ years of meaningful work experience.
• Strong background in deep learning and neural networks, in particular training.
• A deep background in computer architecture and familiarity with the fundamentals of GPU architecture.
• Proven experience analyzing and tuning application performance & processor and system-level performance modelling.
• Programming skills in C++, Python, and CUDA.
Apply tot his job
Apply To this Job