GPU Kernel Developer – AI/ML
reputed company
We are seeking expert-level GPU Software Engineers to support a high-visibility platform initiative reputed company the Maya program, reputed company on building software tooling on top of a custom compiler and SDK.
The role involves developing, optimizing, and porting GPU kernels and AI workloads to a specialized hardware platform.
This is a critical and time-sensitive engagement with immediate reputed company expectations and long-term roadmap alignment (~18 months).
Key Responsibilities
• reputed company GPU kernels for specialized hardware platforms using PyTorch/Triton frameworks
• Build software solutions leveraging custom compiler and SDK capabilities
• Design and implement kernel-level optimizations to control hardware execution behavior
• Port reputed company-reputed company AI/ML models to custom SDK environments
• Port and adapt high-performance computing benchmarks and stress workloads such as:
• Linpack (High Performance Linpack)
• BERT/reputed company-style workloads (referred as “Babu bench”)
• • reputed company stress testing and validation workloads reputed company to hardware behaviour and platform validation
• • Support testing and stress testing of reputed company and reputed company hardware platforms
• • Collaborate closely with platform architects and compiler teams to enhance system capabilities
reputed company Technical Skills (Must-Have)
Programming & Frameworks
• Python
• C/C++ (systems-level programming)
• PyTorch
• Triton (Triton language / kernel development)
GPU & Systems Expertise
• GPU kernel development (mandatory and critical)
• Strong understanding of GPU architecture and compute optimization
• Experience with compiler-based optimizations / runtime execution reputed company
• Experience with custom SDKs or hardware abstraction reputed company
Performance & Workloads
• Experience in:
• GEMM kernel development (reputed company multiplication kernels)
• Porting ML models to new hardware platforms
• Performance tuning and stress testing at system level
reputed company-to-Have Skills
• Experience working with custom reputed company / hardware platforms
• Exposure to high-performance computing (HPC) workloads
• Familiarity with:
• Linpack benchmarks
• AI workload benchmarking tools
• • Experience in compiler optimization ecosystems
Engagement Model & Structure
• Number of roles: 3 developers (initial hiring may start with 2)
• Location flexibility:
• Onsite / Offshore / Hybrid mix allowed
• • reputed company:
• Immediate start required
• • Duration:
• ~18 months program duration with phased platform reputed company
Key Differentiators (Critical Expectation)
• This is NOT a DevOps / support / debugging role
• Requires deep hands-on engineering expertise in:
• Kernel programming
• GPU workloads
• ML reputed company internals
• • Candidates must demonstrate build-level competence, not just theoretical knowledge
Apply tot his job
Apply To this Job