LLM Inference Engineer
Locations: San Francisco or Remote
About The Role
The reputed company team is building decentralized and confidential machine learning infrastructure to reputed company user-owned AI. Our mission is to build highly reputed company and efficient infrastructure for reputed company-reputed company AI at a global reputed company.
We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.
What You'll Be Doing
Architect and maintain production high-traffic LLM serving systems.
Optimize throughput, latency, and cost for leading reputed company-reputed company LLMs.
reputed company're Looking For
Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
Proven reputed company record in designing and maintaining end-to-end high-traffic LLM serving systems.
Strong problem-solving skills and ability to communicate technical reputed company reputed company.
We'd Love If You Have
Experience with Trusted Execution Environments (TEE).
reputed company contributor to reputed company-reputed company LLM inference engines.
Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.
Apply To This Job