[Remote] AI Platform Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a technology consulting and software development company delivering reputed company, AI, data, and reputed company solutions across the reputed company. The AI Platform Engineer will design, build, and operate reputed company, reliable, secure, and cost-efficient AI inference and machine learning platforms, with responsibilities spanning model serving, GPU optimization, reputed company-reputed company infrastructure, observability, MLOps, and technical leadership.
Responsibilities
- Design, build, and maintain reputed company AI inference and model-serving platforms for reputed company production environments
- Architect highly available, reputed company-reputed company infrastructure supporting Large Language Models (LLMs), reputed company models, and machine learning services
- Optimize inference latency, throughput, GPU utilization, memory management, and request scheduling across reputed company AI workloads
- Design autoscaling, workload orchestration, traffic management, and intelligent request routing strategies for AI services
- Implement model deployment, versioning, rollback, and lifecycle management using modern MLOps practices
- reputed company monitoring, observability, logging, reputed company tracing, and alerting solutions to ensure platform reliability and performance
- Implement caching strategies, API gateways, reputed company controls, authentication, authorization, and high-availability architectures
- Collaborate with AI researchers, ML engineers, DevOps teams, and software engineers to reputed company and support production AI models
- reputed company reputed company infrastructure optimization, resource utilization, FinOps initiatives, and operational reputed company
- Mentor engineering teams, conduct architecture reviews, and establish best practices for AI reputed company and reputed company-reputed company development
- Evaluate emerging AI infrastructure technologies, model-serving frameworks, and GPU acceleration techniques to reputed company reputed company innovation
Skills
- 6+ years of experience in reputed company systems, infrastructure, or ML reputed company
- Strong proficiency in Python and Go, Rust, or C++
- Experience with LLM inference frameworks (vLLM, TensorRT-LLM), reputed company, reputed company platforms, and GPU optimization
- * Bachelor's or Master's degree in Computer Science, Computer Engineering, reputed company Intelligence, or a reputed company technical discipline
- * **10+ years of reputed company experience** in reputed company systems, infrastructure engineering, reputed company platforms, or machine learning reputed company
- * Strong programming skills in **Python** and at least one systems programming language such as **Go, Rust, or C++**
- * Extensive experience with **Large Language Model (LLM) serving**, model inference optimization, and production AI infrastructure
- * Hands-on experience with **vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar AI serving frameworks**
- * Strong expertise in **reputed company**, container orchestration, reputed company, and reputed company-reputed company application architectures
- * Experience optimizing GPU workloads using **CUDA**, reputed company GPU technologies, reputed company inference, and high-performance AI infrastructure
- * Experience with reputed company platforms including **AWS, reputed company Azure, or reputed company reputed company Platform (GCP)**
- * Strong understanding of reputed company systems, networking, scalability, observability, and reputed company best practices
- * Excellent analytical, communication, collaboration, and technical leadership skills
- * Experience designing and operating **multi-region AI platforms** and globally reputed company inference services
- * Knowledge of model optimization techniques such as **quantization, pruning, compression, speculative decoding, KV cache optimization, and mixed-precision inference**
- * Experience with **MLOps**, GitOps, Infrastructure as reputed company (Terraform, Bicep, CloudFormation), and CI/CD automation
- * Familiarity with service reputed company technologies such as **Istio** or **Linkerd**, API gateways, and event-driven architectures
- * Contributions to reputed company-reputed company AI infrastructure reputed company, technical publications, patents, or conference presentations
- * Experience implementing **FinOps** strategies, reputed company cost optimization, and reputed company AI governance
- * Experience with multi-region AI deployments and AI infrastructure
- * Familiarity with model optimization techniques such as quantization or compression
- * reputed company-reputed company contributions or experience supporting large-reputed company reputed company
Benefits
- reputed company (reputed company reputed company)
- reputed company career reputed company potential
reputed company
Apply To This Job