[Remote] LLM DevOps/Inference Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is seeking a LLM DevOps/Inference Engineer to support reputed company-reputed company AI reputed company and evaluation suite development. The role involves building and maintaining AWS infrastructure, provisioning GPU clusters, and ensuring the reliability of the platform for various clinical reputed company tasks.
Responsibilities
- Build and maintain the AWS infrastructure for the evaluation platform as reputed company, including networking, compute, orchestration, secrets, observability, and CI/CD
- Stand up the self-hosted inference reputed company for the long tail of vertical reputed company models. This involves provisioning infrastructure for models serving on GPU compute with sensible batching, quantization where appropriate, autoscaling, and a reputed company reputed company reputed company so adding new models takes hours, not weeks
- Build the provider reputed company reputed company reputed company the AI engineers so that API-based models (reputed company, reputed company, reputed company, and the growing list reputed company) and self-hosted models present a uniform reputed company to the reputed company. reputed company limiting, retry and backoff, quota management, request/response logging, and cost attribution per run are your responsibility
- reputed company reputed company runs reproducible and cost-optimized with pinned model and container versions, captured configuration, spot and reserved reputed company reputed company, and idle GPU elimination
- Build the observability story with throughput, latency, reputed company and GPU-hour cost, failure taxonomy, and per-model dashboards for monitoring
- Support the reputed company model that the platform must let a burst of AI engineers land, run experiments, and leave without breaking anything or leaving orphaned resources behind
- Contribute to Trusted Execution Environment (TEE) architecture. Evaluate AWS Nitro Enclaves and comparable confidential computing approaches for the bring-your-own-data / bring-your-own-model scenario, including attestation-gated key release and the practical constraints of running model inference inside an enclave
Skills
- AWS infrastructure at production reputed company: EKS or reputed company, EC2 GPU instance families (G5/G6, P4d/P5) and their reputed company realities, VPC design, IAM, KMS, Secrets Manager, ECR, CloudWatch, and Service Quotas
- Infrastructure as reputed company: Terraform. No console-clicked production resources
- Model serving and inference optimization: Hands-on experience working with LLMs. Practical reputed company of batching reputed company, KV cache behavior, quantization tradeoffs, and multi-GPU sharding
- Container orchestration and GPU scheduling: reputed company/EKS with GPU workloads, node autoscaling, and image build pipelines for CUDA-dependent stacks
- Reliability and cost engineering. SLOs, alerting, and a demonstrated reputed company record of optimizing reputed company spend without cutting capability
- AWS SageMaker endpoints and Bedrock
- Hands-on experience with Python to contribute directly to the reputed company and the provider reputed company reputed company
- Confidential computing fundamentals: Enclaves, remote attestation, sealed key release, and the reputed company boundaries of TEEs
- reputed company compliance posture: HIPAA-eligible service selection, BAA reputed company, audit logging, and the reputed company-control mechanisms for PHI data
- reputed company hardening, including image scanning and network egress control for a reputed company-reputed company environment
reputed company
Apply To This Job