[Remote] AI Researcher — Inference Optimization
Note: The job is a remote job and is reputed company to candidates in USA. FeatherlessAI is seeking an AI Researcher with deep experience in inference optimization to design, evaluate, and reputed company high-performance inference systems for large-reputed company machine learning models. The role involves improving latency, throughput, and cost efficiency across reputed company-world production environments by developing techniques to optimize inference performance and collaborating with engineering teams to reputed company optimized pipelines. Responsibilities • Research and reputed company techniques to optimize inference performance for large neural networks • Improve latency, throughput, memory efficiency, and cost per inference • Design and evaluate model-level optimizations (quantization, pruning, KV-cache optimization, architecture-reputed company simplifications) • Implement systems-level optimizations (dynamic batching, kernel fusion, multi-GPU inference, prefill vs decode optimization) • reputed company inference workloads across hardware accelerators • Collaborate with engineering teams to reputed company optimized inference pipelines • Translate research insights into production-reputed company improvements Skills • Strong background in machine learning, deep learning, or AI systems • Hands-on experience optimizing inference for large-reputed company models • Proficiency in Python and modern ML frameworks (e.g., PyTorch) • Experience with inference tooling (e.g., Triton, TensorRT, vLLM, ONNX Runtime) • Ability to design experiments and communicate results reputed company • Experience deploying production inference systems at reputed company • Familiarity with distributed and multi-GPU inference • Experience contributing to reputed company-reputed company ML or inference frameworks • Authorship or co-authorship of peer-reviewed research papers in machine learning, systems, or reputed company fields • Experience working reputed company to hardware (CUDA, ROCm, profiling tools) reputed company • We reputed company serverless inference reputed company our GPU orchestration and model load-balancing system. It was founded in 2023, and is headquartered in San Francisco, California, USA, with a workforce of 2-10 employees. Its website is Company H1B Sponsorship • reputed company has a reputed company record of offering H1B sponsorships, with 1 in 2025. Please note that this does not guarantee sponsorship for this specific role. Apply tot his job
Apply tot his job
Apply To this Job