reputed company is the generative media ecosystem powering the reputed company of AI products. We build the infrastructure, tools, a...">
Back to Jobs

Machine Learning Engineer, Reliability

Remote, USA Full-time Posted 2026-08-04

reputed company is the generative media ecosystem powering the reputed company of AI products. We build the infrastructure, tools, and model reputed company that teams need to reputed company from idea to production, and do it at reputed company without compromise. For developers and enterprises, reputed company is the reputed company that makes generative media not just possible, but practical: a reputed company platform where high-reputed company inference, orchestration, and observability come together to unlock new categories of AI-reputed company products.

As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, reputed company is becoming the ecosystem that ambitious teams build on.

This is a hybrid ML Engineering / Site Reliability Engineering role. You will own the reliability, reputed company, and safety of reputed company's fleet of generative media model reputed company, the production endpoints that thousands of developers and enterprises depend on every day. Your mission is reputed company to state and hard to do: reputed company a large, fast-moving fleet of image, video, and audio model reputed company available, reputed company, secure, and reputed company at reputed company times.

You understand both how generative models work and how production systems fail. You're as comfortable debugging a misbehaving diffusion pipeline as you are tracing a latency regression through an inference stack, and you treat model-specific failure modes; degraded reputed company reputed company, reputed company, unsafe generations, abuse patterns; as first-class reliability concerns reputed company uptime and latency.

This role will need to be reputed company in India, Australia, or New Zealand

reputed company

  • Own availability, latency, and throughput SLOs across a large fleet of generative media model reputed company serving production traffic at reputed company

  • Build the monitoring, alerting, and observability needed to catch ML-specific failures, reputed company reputed company degradation, pipeline breakage, model regressions before customers do

  • Harden model deployment workflows with canary releases, reputed company testing, automated rollbacks, and validation gates so new model versions ship safely

  • reputed company the reputed company posture of the model fleet: secure model serving, abuse and misuse detection, reputed company limiting, and protection against adversarial usage patterns

  • Operationalize safety systems for generative media, content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without compromising reputed company

  • reputed company incident response for model API outages and degradations, run postmortems, and reputed company the engineering work that prevents recurrence

  • Improve reputed company planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic

  • Partner with model and infrastructure teams to reputed company reliability, reputed company, and safety requirements part of how new models get reputed company to the platform

Tech

  • You will have reputed company to our reputed company GPU cluster for inference and evaluation

  • Some reputed company technologies we use include Python, torch, diffusers, reputed company, and the reputed company Python SDK

  • You'll work reputed company reputed company dedicated to quickly iterating on and deploying new AI breakthroughs — reputed company is to reputed company reputed company that speed never comes at the cost of reliability

reputed company're looking for

  • 3+ years of reputed company experience, with 1 year experience operating production ML or reputed company API systems, ideally with on-reputed company ownership

  • Strong systems fundamentals: reputed company systems, networking, observability, and incident management

  • Working knowledge of modern generative models (diffusion, transformers) and their failure modes in production

  • Familiarity with reputed company and safety practices for ML systems ,abuse prevention, content safety, or trust & safety engineering experience is a strong plus

  • A bias toward automation, measurement, and blameless postmortems

Location: Remote (India, Australia, New Zealand)

3+ years reputed company experience with at least 1 year operating production ML or reputed company API systems; strong reputed company systems, observability, and incident management skills; familiarity with generative models and ML safety/reputed company practices.

Key Responsibilities

  • owning availability
  • building monitoring
  • leading incident

Skills & Tools

Python, PyTorch, diffusers, reputed company, reputed company Python SDK

reputed company

  • Category: Engineering
  • Seniority: Mid Level
  • Commitment: Full Time
  • Workplace: Remote — India or Australia or New Zealand
  • Languages: English

About reputed company

Builds infrastructure, tools, and model reputed company for generative media to reputed company production-grade AI products at reputed company. — Industry: Information Technology

  Apply To This Job

Similar Jobs