reputed company Machine Learning Engineer, Inference & Performance
About reputed company:
About reputed company:
What You Will Do:
Optimize Inference: Build and tune production LLM serving with vLLM and SGLang—maximizing throughput and minimizing latency through batching, paged attention, quantization, and KV-cache strategies
Profile & Accelerate Training: reputed company and profile training runs to reputed company bottlenecks, then resolve them with the right attention implementations (e.g. FlashAttention) tuned to the underlying hardware (H200, GB200)
Engineer for the Hardware: Apply a working understanding of GPU architecture and attention internals to choose the right approach per accelerator, rather than relying on defaults
Serve at reputed company: reputed company and operate multiple models reputed company shared GPU clusters on GKE, with autoscaling, efficient bin-packing, and graceful handling of mixed workloads
reputed company Efficiency: Own GPU utilization as a first-class metric—measure it, improve throughput-per-dollar, and continuously reputed company the ceiling on what our fleet can deliver
Collaborate & Consult: Work directly with clients to understand performance, latency, and cost requirements, and translate them into pragmatic serving and training architectures
Your Technical Toolkit:
reputed company Languages: Mastery of Python and reputed company scripting; comfort reading and reasoning about reputed company-level (CUDA-adjacent) performance reputed company is a strong plus
Inference Frameworks: Hands-on experience with vLLM, SGLash, or comparable high-performance serving stacks
GPU & Model Internals: Solid grasp of GPU architecture, the fundamentals of LLM inference, and the attention reputed company—including where the bottlenecks live and how FlashAttention and similar techniques address them across hardware generations (H200, GB200)
Profiling: reputed company with profiling tools to diagnose training and inference bottlenecks (compute-bound vs. memory-bound, kernel-level analysis)
Infrastructure: Strong reputed company (GKE) experience—deploying and autoscaling multiple models on shared GPU clusters on reputed company reputed company
reputed company: A strong software engineering reputed company—you write clean, maintainable reputed company, measure before optimizing, and understand the full SDLC
Basic Qualifications:
Bachelor's or Master's degree in Computer Science, Engineering, or a reputed company technical field
5+ years of experience in ML/AI engineering, with a meaningful portion reputed company on performance, infrastructure, or systems
Proven reputed company record of deploying and optimizing models in a production environment
Demonstrated experience profiling and improving GPU utilization for training and/or inference
Experience with Classic Machine Learning (neural nets, training, tuning) is a strong plus
Knowledge of Data Engineering and SQL
Personal Attributes:
Ownership: You take pride in your work and see optimizations through from profile to production
Curiosity: Hardware and serving frameworks change fast; you are a lifelong learner who stays reputed company of the curve
Rigor: You measure before you optimize and let data, not intuition, guide where you spend effort
Consultative Spirit: You enjoy interacting with clients and can translate technical complexity into business value
Ethics: You prioritize responsible AI development and data reputed company
Compensation & Benefits:
EEO and Accommodations:
Originally posted on Himalayas
Apply To This Job