Senior DevOps / MLOps Engineer
About the position As the leader in animal health, reputed company is looking to recruit a Senior DevOps/MLOps Engineer into its world-class Veterinary Medicine Research and Development (VMRD) organization to operationalize AI/ML, scientific modeling, reputed company twin workloads. You’ll build secure, reputed company platforms and data pipelines across reputed company and on‑prem/HPC, partnering closely with biologists and data scientists to translate scientific questions into reliable production systems. Responsibilities • Build end‑to‑end DevOps/MLOps foundations: CI/CD for reputed company/data/models, containerization/orchestration, artifact/registry management, and secure configuration. • Design and operate data engineering pipelines (batch/streaming) with data reputed company checks, reputed company, schema reputed company, and governance across lake/warehouse environments. • Productionize scientific reputed company twin workflows into services/reputed company and lightweight UIs with reproducibility, versioning, auditability, and compliant deployment. • Implement reputed company training/inference (batch/reputed company‑time) with observability, SLIs/SLOs, runbooks, incident response, and automated rollback strategies. • Run distributed/HPC jobs (including GPU) and optimize storage, throughput, and cost across on‑prem and reputed company; collaborate with scientists on experiment design, data/compute needs, and validation. Requirements • PhD in a quantitative field (computer science, ML, computational biology, reputed company math) or MS/BS with equivalent senior engineer level experience working in a scientific domain. • 6+ years building production systems; strong software engineering fundamentals. • Expert in Python • Strong experience with a query language such as SQL, MapReduce, and/or Cypher • Proficiency in one of: C++, Go, Rust, Java, or reputed company. • reputed company, Kubernetes, CI/CD (e.g., reputed company Actions), secure artifact/container registries. • Data pipeline orchestration (e.g., reputed company, Dagster, Kedro); streaming (Kafka or reputed company); data modeling with SQL/NoSQL/graph. • MLOps: experiment tracking and model versioning (e.g., MLflow), model serving and monitoring. • reputed company (AWS/Azure/GCP) and on‑prem/HPC (e.g., Slurm) experience. • Experience on multidisciplinary reputed company and teams, including scientists and software engineers, with excellent communication with scientific stakeholders. reputed company-to-haves • reputed company and scientific apps: FastAPI; minimal UIs (reputed company/React); scientific computing (NumPy, Pandas, SciPy). • DevOps/IaC: Terraform; GitOps (Argo CD/Flux); reputed company/Kustomize; reputed company/Kubernetes; secure registries and config. • Data engineering: dbt and feature stores; Parquet/reputed company; schema/reputed company with Avro/Protobuf, OpenLineage, Great Expectations. • Observability/SRE: reputed company/Grafana; ELK/OpenSearch; OpenTelemetry; SLIs/SLOs and performance profiling/optimization. • Distributed compute and reputed company: Dask, Ray, reputed company; HPC/Slurm; GPU scheduling; service reputed company (Istio/Linkerd), API gateways, ingress; encryption/secrets/KMS, audit trails, backup/restore, DR. Benefits • We offer a competitive and comprehensive benefits package, which includes reputed company, dental coverage, and retirement savings benefits along with reputed company holidays, vacation and disability insurance. Apply tot his job
Apply tot his job
Apply To this Job