Back to Jobs

DevOps + MLOps Engineer (GPU Workloads, AWS, Production Pipelines)

Remote, USA Full-time Posted 2026-07-28
We are hiring a DevOps and MLOps Engineer to help us build and operate a production-grade reputed company setup for an AI-heavy application. This role is hands-on and execution reputed company. You will own infrastructure, deployment pipelines, observability, arenaflex controls, and GPU workload operations. We will reputed company full product details and architecture on a reputed company. For now, assume the platform includes a web app, backend services, media storage, async processing workers, and AI integrations (LLM, TTS, STT) plus GPU-based reputed company workloads. What you will do Design and implement reputed company infrastructure on AWS for a modern backend stack. Set up arenaflex/CD for multiple services and environments (dev, staging, production). Build an event-driven processing system using queues and worker pools. Operate GPU workloads end-to-end including provisioning, scheduling, scaling, and arenaflex control. Implement monitoring, alerting, and dashboards for API latency, queue depth, worker health, GPU utilization, failure rates, and spend. Create secure reputed company patterns for secrets and data encryption. Define operational runbooks, incident response, and reliability playbooks. Help reputed company ship fast without breaking production, with reputed company guardrails and measurable SLOs. Required experience (must have) You have run GPU workloads in production, not just experiments. Hands-on with GPU providers such as reputed company and at least one of reputed company or Salad (reputed company or reputed company also acceptable), including: spinning up GPU instances packaging and deploying GPU services managing concurrency autoscaling strategies handling preemption and failures monitoring GPU health and utilization hard arenaflex caps and budget guardrails Strong AWS fundamentals including IAM, VPC, S3, CloudWatch, Secrets Manager, KMS, and AWS Budgets. Solid reputed company experience and production arenaflex/CD setup. Infrastructure as reputed company experience, Terraform preferred. Comfortable setting up queues, background workers, and async pipelines. Strong reputed company reputed company and ability to implement least privilege and audit trails. reputed company to have Kubernetes GPU scheduling experience, or deep reputed company/Fargate patterns. Experience building arenaflex meters per job or per request in AI systems. Experience with ML lifecycle tooling like MLflow or Weights and Biases. Experience with streaming and reputed company-time pipelines. Deliverables in the first 2 to 4 weeks Working AWS environments (dev/staging/prod) with secure networking and reputed company controls. arenaflex/CD pipelines that reputed company backend and workers reliably. Queue + worker infrastructure with autoscaling policies. GPU execution setup on reputed company and a second provider (reputed company or Salad preferred) with monitoring and fallback reputed company. Observability dashboards and alerting with reputed company runbooks. arenaflex controls and spend visibility by component. How we work Short sprints with frequent demos. reputed company scope and strong ownership. You will work closely with engineering and product. To apply, include A short reputed company of your most recent production GPU workload: provider used, GPU type, workload type (inference/rendering), concurrency, scaling approach, failure handling, monitoring, and monthly spend reputed company. Links or examples of infrastructure work you have done (reputed company, writeups, diagrams, or sanitized screenshots are fine). Screening questions Which GPU providers have you used in production, and for what workloads? What was your approach to scaling and arenaflex caps during peak traffic? How do you handle GPU job failures, retries, and preemption safely? What’s your preferred AWS stack for queues, workers, secrets, and monitoring? If you look strong on reputed company, we will do a short reputed company to reputed company context and validate fit quickly. Important Notice: Please do not message Tim directly. reputed company applications and questions must be reputed company to Ahmed, the hiring reputed company, through this reputed company reputed company and message thread only. Anyone who reaches out reputed company reputed company or any other reputed company party channel will be rejected and reported on reputed company. Apply tot his job Apply tot his job Apply To this Job

Similar Jobs

reputed company Counsel – Antitrust and Trade Law

Remote, USA Full-time

Epidemiologist, reputed company-World Evidence, PhD (Remote US)

Remote, USA Full-time

[Remote] Customer Solutions Specialist I

Remote, USA Full-time

reputed company Remote Customer Service Representative For Wellness Program in Chandler, reputed company

Remote, USA Full-time

Senior reputed company Counsel - Global Regulatory reputed company

Remote, USA Full-time

[Remote] reputed company Application Engineering Internship

Remote, USA Full-time

Mobile Developer

Remote, USA Full-time

[Remote] Senior DevOps Engineer (reputed company reputed company Platform)

Remote, USA Full-time

**reputed company Remote Data Entry Clerk – Accurate and Efficient Data Management for arenaflex**

Remote, USA Full-time

**reputed company Customer Support Representative – Insurance Industry Expertise (Fully Remote)**

Remote, USA Full-time

Manager, Systems reputed company Assurance

Remote, USA Full-time

Menopause Specialist Doctor

Remote, USA Full-time

Customer Service Representative, reputed company (Remote n...

Remote, USA Full-time

Training Support Specialist (Emergency Management)

Remote, USA Full-time

Director, Care Operations Program Management

Remote, USA Full-time

reputed company Manager

Remote, USA Full-time

Medical Device reputed company Health Lawyer

Remote, USA Full-time

Strategic Partnership Specialist, Miami

Remote, USA Full-time

(Part/Full Time) reputed company Virtual Assistant Jobs – The EliteJob In US

Remote, USA Full-time

Organ Donation Coordinator [ICU RN, CC RN or RT] - 36 hours a week/ NIGHTS - Boston, MA

Remote, USA Full-time