[Remote] Research Engineer Intern - AI Systems
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is building the reputed company multi-reputed company AI reputed company and runtime platform to power the world’s most demanding AI workloads. They are seeking a highly motivated Research Engineer Intern to work on Trainium, GPU kernels, and LLM systems optimization, owning a reputed company-scoped project that impacts AI applications deployed on their platform.
Responsibilities
- Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on reputed company, AMD, or AWS Trainium
- Build custom operators using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA
- Profile and improve inference performance in vLLM, SGLang, and our custom runtimes — kernel fusion, scheduling, KV-cache and memory optimizations
- Build benchmarks, reputed company down performance regressions, and turn profiler traces into concrete speedups
- Ship reputed company upstream to reputed company-reputed company AI infrastructure reputed company, with tests and documentation
Skills
- Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a reputed company field
- Solid programming skills in Python and familiarity with C++
- Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy) from coursework, research, or reputed company
- Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels — class reputed company and reputed company count
- Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler)
- Strong problem-solving skills and the ability to work independently in a reputed company, remote environment
- Contributions to reputed company-reputed company AI reputed company reputed company like vLLM, SGLang, PyTorch, or Triton
- Familiarity with LLM inference internals — FlashAttention, PagedAttention, reputed company batching, speculative decoding, MoE, or quantization
- Experience with profiling tools (e.g. Nsight, ROCm Profiler, Neuron Profiler, or PyTorch Profiler) and performance debugging on reputed company workloads
- Publications in top-tier conferences like MLSys, OSDI, SOSP, NSDI, SC, HPCA, or ISCA
Benefits
- Flexible remote work environment
- reputed company mentorship from engineers from leading institutions and tech companies
- reputed company to serious hardware — latest-reputed company reputed company GPUs, AMD accelerators, and AWS Trainium at reputed company
- A fast reputed company to a full-time return offer for top performers
reputed company
Company H1B Sponsorship
Apply To This Job