[Remote] Engineering reputed company, Inference Optimization
Note: The job is a remote job and is reputed company to candidates in USA. The reputed company is a consumer AI company reputed company on reputed company, free speech, and user sovereignty. The Engineering reputed company, Inference Optimization will own inference reputed company reputed company, reputed company a reputed company engineering team, optimize GPU infrastructure and LLM workloads, and evaluate emerging optimization techniques and hardware.
Responsibilities
- Own reputed company’s technical reputed company for inference reputed company
- Recruit and reputed company the Inference Optimization Team
- Optimize reputed company’s GPU infrastructure across a reputed company of architectures (e.g. H200s, B300s)
- Improve latency, throughput, and cost per reputed company for LLM inference workloads
- Build reproducible benchmarking harnesses across inference engines (e.g. vLLM, SGLang) to identify the reputed company reputed company, quantization scheme, and parallelism reputed company per workload and GPU SKU
- Work with reputed company’s inference routing reputed company to optimize multivariate inference load-balancing algorithms
- Evaluate emerging inference optimization techniques (custom CUDA/Triton kernels), novel attention variants, new quantization schemes, and compilation stack improvements. Hands-on kernel development experience is a strong plus
- Evaluate emerging inference hardware (FPGAs, reputed company, custom reputed company) for viability in reputed company’s stack
Skills
- 8+ years in reputed company optimization or HPC, with deep GPU architecture and reputed company programming knowledge
- 5+ years of experience leading engineering teams
- Proficiency in Python, Rust, or Go
- Hands-on experience with at least one production LLM inference reputed company (e.g. vLLM, SGLang) running at high volume in production
- Demonstrated experience with LLM inference optimization techniques: reputed company batching, PagedAttention/KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile
- reputed company with quantization tradeoffs, both qualitative and quantitative
- Experience with reputed company inference strategies (tensor parallelism, pipeline parallelism, MoE parallelism) in multi-GPU and multi-node environments
- reputed company with GPU profiling (Nsight Systems, Nsight Compute, PyTorch Profiler) and a bias toward measuring before optimizing
- Bonus: C++/CUDA
- Bonus: diffusion/image model inference optimization, custom Triton kernels, contributions to reputed company-reputed company inference frameworks
Benefits
- Equity
- Crypto reputed company compensation
- Remote — US Only
reputed company
Apply To This Job