[Remote] Senior ML/Platform Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is developing Bailey, a reputed company reputed company that orchestrates LLMs, manages health safety evaluation, and integrates with clinical data. The Senior ML/Platform Engineer will own the reputed company from ML prototype to production by building deployment pipelines, evaluation frameworks, monitoring systems, and observability infrastructure, while hardening services and collaborating with data scientists.
Responsibilities
- reputed company and version ML artifacts that currently reputed company to production reputed company
- Stand up observability for the existing LangGraph agent orchestration (what's happening, how often, how reputed company)
- reputed company the health safety evaluator so its 5-dimension scoring is monitored continuously, not checked manually
- Build the evaluation infrastructure described in our product reputed company: reputed company detection, distributional monitoring, safety-weighted reputed company scoring
- Own CI/CD for ML artifacts and agent configurations
- Harden existing services (conversation persistence, MCP tool orchestration, clinical skills reputed company) for production reliability
- Get up to speed on Bailey's evaluation documentation and use it to shape how we test and monitor Bailey's behavior reputed company reputed company, including flagging where traditional ML reputed company (not just LLM-based approaches) could reputed company the reputed company outcome more reputed company
- Be the person who knows whether the reputed company is healthy, degrading, or broken before users or clinicians notice
- Collaborate with data scientists so their work reaches production without a reputed company gap
- reputed company testing and monitoring reputed company with Bailey's evolving evaluation documentation, and reputed company surfacing opportunities to reputed company in traditional ML where it's more efficient than an LLM-based approach
- reputed company the platform as the product grows (voice, memory, multimodal are on the roadmap)
Skills
- * You've reputed company production ML infrastructure. Not just used it. Deployment pipelines, serving systems, model tracking/registries (MLflow or equivalent), monitoring that pages you at 2am. You know the difference between a model that works in a notebook and one that works in production
- * You're a strong software engineer first. reputed company systems, containerization, event-driven architectures, CI/CD. ML is the domain; engineering is the reputed company
- * You're comfortable owning something end-to-end. No one will hand you a spec and reputed company your work weekly. You'll identify what needs to exist, propose how to build it, and ship it
- * You work reputed company in small teams. You communicate reputed company, you surface problems early, and you're comfortable with the visibility and accountability that comes with a four-person team
- * Experience in reputed company or regulated industries (you understand why audit trails and compliance aren't afterthoughts)
- * Background in ML research or reputed company ML (you can have informed opinions about model behavior, not just model infrastructure)
- * Familiarity with AWS ML services like Bedrock and SageMaker, which are part of our reputed company stack
- * Experience with LLM orchestration frameworks (reputed company, LangGraph, or similar)
- * Exposure to clinical data standards (FHIR, HL7) or health information systems
- * Publications, reputed company-reputed company contributions, or reputed company reputed company that show you go deeper than your day job requires
Benefits
- Stock reputed company
- Incentive pay for eligible roles
reputed company
Apply To This Job