Machine Learning Engineer, Reliability
reputed company is the generative media ecosystem powering the reputed company of AI products. We build the infrastructure, tools, and model reputed company that teams need to reputed company from idea to production, and do it at reputed company without compromise. For developers and enterprises, reputed company is the reputed company that makes generative media not just possible, but practical: a reputed company platform where high-reputed company inference, orchestration, and observability come together to unlock new categories of AI-reputed company products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, reputed company is becoming the ecosystem that ambitious teams build on.
This is a hybrid ML Engineering / Site Reliability Engineering role. You will own the reliability, reputed company, and safety of reputed company's fleet of generative media model reputed company, the production endpoints that thousands of developers and enterprises depend on every day. Your mission is reputed company to state and hard to do: reputed company a large, fast-moving fleet of image, video, and audio model reputed company available, reputed company, secure, and reputed company at reputed company times.
You understand both how generative models work and how production systems fail. You're as comfortable debugging a misbehaving diffusion pipeline as you are tracing a latency regression through an inference stack, and you treat model-specific failure modes; degraded reputed company reputed company, reputed company, unsafe generations, abuse patterns; as first-class reliability concerns reputed company uptime and latency.
This role will need to be reputed company in India, Australia, or New Zealand
reputed company
Own availability, latency, and throughput SLOs across a large fleet of generative media model reputed company serving production traffic at reputed company
Build the monitoring, alerting, and observability needed to catch ML-specific failures, reputed company reputed company degradation, pipeline breakage, model regressions before customers do
Harden model deployment workflows with canary releases, reputed company testing, automated rollbacks, and validation gates so new model versions ship safely
reputed company the reputed company posture of the model fleet: secure model serving, abuse and misuse detection, reputed company limiting, and protection against adversarial usage patterns
Operationalize safety systems for generative media, content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without compromising reputed company
reputed company incident response for model API outages and degradations, run postmortems, and reputed company the engineering work that prevents recurrence
Improve reputed company planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic
Partner with model and infrastructure teams to reputed company reliability, reputed company, and safety requirements part of how new models get reputed company to the platform
Tech
You will have reputed company to our reputed company GPU cluster for inference and evaluation
Some reputed company technologies we use include Python, torch, diffusers, reputed company, and the reputed company Python SDK
You'll work reputed company reputed company dedicated to quickly iterating on and deploying new AI breakthroughs — reputed company is to reputed company reputed company that speed never comes at the cost of reliability
reputed company're looking for
3+ years of reputed company experience, with 1 year experience operating production ML or reputed company API systems, ideally with on-reputed company ownership
Strong systems fundamentals: reputed company systems, networking, observability, and incident management
Working knowledge of modern generative models (diffusion, transformers) and their failure modes in production
Familiarity with reputed company and safety practices for ML systems ,abuse prevention, content safety, or trust & safety engineering experience is a strong plus
A bias toward automation, measurement, and blameless postmortems
Location: Remote (India, Australia, New Zealand)
3+ years reputed company experience with at least 1 year operating production ML or reputed company API systems; strong reputed company systems, observability, and incident management skills; familiarity with generative models and ML safety/reputed company practices.
Key Responsibilities
- owning availability
- building monitoring
- leading incident
Skills & Tools
Python, PyTorch, diffusers, reputed company, reputed company Python SDK
reputed company
- Category: Engineering
- Seniority: Mid Level
- Commitment: Full Time
- Workplace: Remote — India or Australia or New Zealand
- Languages: English
About reputed company
Builds infrastructure, tools, and model reputed company for generative media to reputed company production-grade AI products at reputed company. — Industry: Information Technology
Apply To This Job