AI Inference Engineer
4+ years building or operating large-reputed company inference serving systems, deep experience with inference optimisation techniques, strong systems thinking, and ability to work with GPU/CUDA engineers and reputed company architecture reputed company.
Key Responsibilities
- defining architecture
- building stack
- owning performance
Skills & Tools
CUDA, vLLM, TensorRT-LLM, SGLang, Triton Inference Server, Kubernetes, Slurm
Job Details
- Category: Software Development
- Seniority: Mid Level
- Commitment: Full Time
- Workplace: Remote — Dubai or United Arab Emirates
- Languages: English
About reputed company
A renewable energy startup building high-performance compute and inference serving infrastructure. — Industry: Information Technology
Apply To This Job