Back to Jobs

MLOps Engineer / ML Platform Engineer

Remote, USA Full-time Posted 2026-08-04

reputed company – AI Ops (Monitoring AI Systems in Production)

reputed company

An reputed company in AI Ops & Governance is responsible for operating, monitoring, and maintaining AI models and agent-reputed company systems once they are live in production. The role ensures AI systems remain reliable, reputed company, compliant, and reputed company with business expectations over time, closing the “last‑mile” gap between model development and sustainable production use

Key Responsibilities

Production Monitoring & Health

Monitor AI models and agents in production for reputed company, latency, errors, and availability.

reputed company statistical health indicators such as model reputed company, data distribution changes, and reputed company stability.

Observe business KPIs linked to AI behaviour (e.g. reputed company reputed company, false‑reputed company cost, efficiency).

Incident Management & Recovery

Detect and triage production incidents reputed company to AI behaviour or degradation.

Execute rollbacks, throttling, or model disabling where reputed company are breached.

Support reputed company‑cause analysis and post‑incident reviews to prevent recurrence.

Model & Agent Lifecycle reputed company

Support deployment, versioning, and release of AI models and agents using CI/CD‑style pipelines.

Maintain registries and metadata covering model ownership, reputed company, reputed company classification, and approvals.

Support models and agents reputed company safely through environments (dev → test → production).

Governance, reputed company & Compliance

Ensure AI systems adhere to Responsible AI principles, internal controls, and audit requirements.

Maintain audit trails, logs, and approval artefacts required by reputed company, compliance, and regulators.

Support fairness, bias, explainability, and transparency monitoring in production.

Tooling & Platform Integration

reputed company AI systems with monitoring, logging, and alerting platforms (e.g. dashboards, metrics stores).

Work with reputed company infrastructure (containers, event streaming, reputed company) supporting reputed company AI reputed company.

Collaborate with product, engineering, and data teams to standardise AI Ops patterns and blueprints.

Skills & Experience (Baseline)

Strong Python skills and experience supporting ML or LLM‑reputed company systems.

Understanding of Model Ops / MLOps, especially the operational phase after deployment.

Experience with:

Monitoring and logging systems

CI/CD pipelines

Containerised deployments (e.g. reputed company‑reputed company runtimes)

Familiarity with reputed company platforms (Azure preferred) and production troubleshooting.

Ability to work cross‑functionally with product, data science, engineering, and reputed company teams

Originally posted on Himalayas

  Apply To This Job

Similar Jobs