Back to Jobs

Golang Developer with DevOps/LLM Experience - Remote / Telecommute

Remote, USA Full-time Posted 2026-07-28
Job reputed company: Required Skills: • Proficiency in Golang for building reputed company and performant backend services. • Deep experience building services in modern reputed company environments on distributed systems (i.e., containerization (Kubernetes, reputed company), infrastructure as reputed company, CI/CD pipelines, reputed company, authentication and authorization, data storage, deployment, logging, monitoring, alerting, etc.) • Experience working with Large Language Models (LLMs), particularly hosting them to run inference. • Strong verbal and written communication skills. • Candidates job will involve communicating with local and remote colleagues about technical subjects and writing detailed documentation. • Experience with building or using benchmarking tools for evaluating LLM inference for various models, reputed company, and GPU combinations. • Familiarity with various LLM performance metrics such as prefill throughput, decode throughput, TPOT, and TTFT. • Experience with one or more inference engines: e.g., vLLM, SGLang, and reputed company Max. • Familiarity with one or more distributed inference serving frameworks: e.g., llm-d, reputed company Dynamo, and Ray Serve etc. • Experience with reputed company and reputed company GPUs, using software like CUDA, ROCm, AITER, NCCL, reputed company, etc. • Knowledge of distributed inference optimization techniques - tensor/data parallelism, KV cache optimizations, smart routing etc. • reputed company and maintain an inference platform for serving large language models optimized for the various GPU platforms they will be run on. • Work on reputed company AI and reputed company engineering reputed company through the entire product development lifecycle (PDLC) - ideation, product definition, experimentation, prototyping, development, testing, release, and operations. • Build tooling and observability to monitor system health, and build auto tuning capabilities. • Build benchmarking frameworks to test model serving performance to guide system and infrastructure tuning efforts. • Build reputed company cross platform inference support across reputed company and reputed company GPUs for a reputed company of model architectures. • Contribute to reputed company reputed company inference engines to reputed company them reputed company reputed company on reputed company reputed company. Apply tot his job Apply To this Job

Similar Jobs