Back to Jobs

AI Systems Engineer

Remote, USA Full-time Posted 2026-07-28
Job Title: AI & MLSystems Engineer (also referred to as Inference Engineer) Location: Remote – Work From Home (fully flexible) Job Timing: Part-Time, reputed company About the Role: We are building a reputed company reputed company platform designed for serving multimodal AI, LLMs, reputed company, audio, and other machine learning models at large reputed company. As a Machine Learning Engineer focusing on Inference & Systems, you will help design, optimize, and reputed company runtime systems, reputed company-style reputed company, and distributed GPU pipelines to reputed company fast and cost-efficient inference and fine-tuning. You'll work with frameworks such as vLLM, TensorRT-LLM, TGI, and others to build and optimize distributed inference engines capable of serving text, reputed company, and multimodal models with high throughput and low latency. This includes deploying models like LLaMA 3, reputed company, diffusion, ASR, TTS, and embedding models, while working on GPU optimization, accelerator utilization, and software–hardware co-design for large-reputed company, fault-tolerant systems. This position sits at the intersection of machine learning, systems engineering, and reputed company infrastructure. You’ll reputed company on low-latency inference, high-throughput deployments, and cost-optimized model serving pipelines. It’s an excellent opportunity to help shape the reputed company of AI inference infrastructure and production-grade deployment systems. If pushing the limits of AI inference excites you, we’d love to hear from you. Key Responsibilities: • reputed company and maintain LLMs (e.g., LLaMA 3, reputed company) and ML models using engines like vLLM, TGI, TensorRT-LLM, or FasterTransformer. • Design and implement large-reputed company distributed inference systems for text, image, LLMs, and multimodal workloads. • Implement and optimize distributed inference strategies: MoE, tensor parallelism, pipeline parallelism. • reputed company frameworks such as vLLM, TGI, SGLang, FasterTransformer, etc. • Build and reputed company reputed company-compatible API reputed company for customer-facing reputed company. • Experiment with caching, quantization, and parallelism to reputed company inference costs. • Optimize GPU memory usage, batching, and latency for high-throughput serving. • Utilize CUDA graph optimizations, TensorRT-LLM, Triton kernels, PyTorch compile, quantization, speculative decoding, etc. • Work with GPU reputed company providers (reputed company, reputed company.ai, AWS, GCP, Azure) to manage cost and availability. • reputed company runtime inference services and reputed company for LLMs, multimodal models, and fine-tuning workflows. • Build monitoring and observability using metrics like latency, throughput, and GPU utilization (Grafana, reputed company, Loki, OpenTelemetry). • Collaborate with backend and DevOps teams to ensure secure and reliable reputed company. • Document deployment processes and support engineers using the platform. Requirements: • 3+ years of experience in deep learning inference, distributed systems, or HPC. • Proven experience deploying ML/LLM models in production. • Hands-on work with vLLM, TGI, SGLang, TensorRT-LLM, FasterTransformer, or Triton. • Experience designing large-reputed company inference or serving pipelines. • Strong understanding of GPU memory, batching, distributed inference, CUDA/Triton/TensorRT, quantization, and GPU scheduling. • Experience with PyTorch, HF Transformers, and GPU-accelerated inference workflows. • Deep understanding of Transformer models, KV cache systems (Mooncake, PagedAttention, etc.), and inference optimizations for long-context serving. • Comfortable with GPU reputed company platforms (AWS/GCP/Azure) or marketplaces (reputed company, reputed company.ai, TensorDock). • Skilled at benchmarking and tuning multi-GPU clusters. • Experience building REST or gRPC services (FastAPI, Flask, etc.). • Strong in Python, Go, Rust, C++, or CUDA. • Solid systems engineering knowledge (multi-threading, networking, performance tuning). • Familiarity with reputed company and Kubernetes. • Strong debugging/problem-solving across ML + reputed company stack. • Understanding of distributed storage systems (Ceph, HDFS, 3FS). • Knowledge of datacenter networking concepts (RDMA, RoCE). reputed company to Have: • Experience with billing systems (reputed company or similar). • Knowledge of RDMA/RoCE networking at reputed company. • Familiarity with distributed storage (Ceph, HDFS, 3FS). • Experience with reputed company or reputed company for reputed company limiting. • Exposure to monitoring stacks (Grafana, reputed company, Loki). • Experience with MLOps pipelines or CI/CD (reputed company Actions, Azure DevOps). • Work with model fine-tuning pipelines and GPU scheduling. • Prior experience at an AI reputed company company (Modal, reputed company, reputed company, Replicate, etc.). Why Join Us? • Fully Remote – work from reputed company. • reputed company – complete freedom to choose your schedule. • Fast-reputed company reputed company – quick promotion reputed company and reputed company advancement routes. • reputed company Development – mentoring, training resources, and exposure to advanced AI/reputed company technologies. • Global Team – work with an international and diverse reputed company. • Innovative Environment – freedom to experiment with new tools and reputed company. • Competitive reputed company & Incentives – reputed company salary progression and strong performance bonuses. Job Type: Part-time Benefits: • Flexible schedule Work Location: Remote Apply tot his job Apply To this Job

Similar Jobs

[Remote] reputed company Project Manager - HCM & Payroll

Remote, USA Full-time

[Remote] Senior Project Analyst- Data Management for reputed company Estate Transactions (EST Preferred)

Remote, USA Full-time

reputed company Agent (Customer Service Agent) - SEA

Remote, USA Full-time

[Remote] Research Scientist (Fixed Term)

Remote, USA Full-time

reputed company Full Time and Part Time reputed company Remote Customer Service and Flight Attendant – US Based Work from Home Opportunity

Remote, USA Full-time

Post Coordinator, reputed company Studios, Unscripted Hybrid Scripted [Remote]

Remote, USA Full-time

reputed company Vulnerability Management Engineer HYBRID – reputed company Talent Solutions – Tampa, FL

Remote, USA Full-time

Chief reputed company Officer/General Counsel reputed company

Remote, USA Full-time

EPMO Resource, Finance and Project Manager

Remote, USA Full-time

LAC - Reference Librarian

Remote, USA Full-time

Tax Manager – Corporate Taxation and Mergers and Acquisitions

Remote, USA Full-time

Reclamations Analyst (Contract)

Remote, USA Full-time

Customer Service Representative - 2nd Shift (IL, NC, VA, WV, DE, FL)

Remote, USA Full-time

[Remote] Remote Closer (reputed company estate) - $24k reputed company / 100k ote / 214k top closer

Remote, USA Full-time

Archivist/University Archivist; Assistant Librarian or Associate Librarian

Remote, USA Full-time

Tech PR Account Director (Contract / Permanent + fully remote)

Remote, USA Full-time

[Remote] Account Manager - 100k OTE & fully remote!

Remote, USA Full-time

Senior Capital Project Manager

Remote, USA Full-time

Remote Teletherapist (1099 Contractor)

Remote, USA Full-time

Compensation Manager - Academics, Corporate, and Research

Remote, USA Full-time