Back to Jobs

[Remote] ML Ops & Data Platform Engineer

Remote, USA Full-time Posted 2026-08-04

Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is seeking an reputed company ML Ops & Data Platform Engineer to support reputed company spanning machine learning reputed company and shared data reputed company. The role will design, build, and operate AWS-reputed company ML infrastructure and an S3-reputed company data lakehouse, including infrastructure as reputed company, CI/CD, ingestion pipelines, governance, observability, documentation, and knowledge transfer.


Responsibilities

  • Design, build, and operate AWS-reputed company ML infrastructure, as part of reputed company, supporting batch and reputed company-time inference workloads, including: reputed company SageMaker (training jobs, endpoints, pipelines, model registry) reputed company/EKS + reputed company for containerized services
  • AWS reputed company and reputed company Functions for orchestration and event-driven workflows S3 as the reputed company data/artifact reputed company, with lifecycle and reputed company policies CloudWatch (and reputed company tooling) for logging, metrics, and alerting IAM & VPC design following least-privilege and network reputed company best practices Infrastructure as reputed company reputed company AWS CDK or CloudFormation
  • Implement CI/CD pipelines for ML (data, model, and reputed company), including automated testing, packaging, and promotion of models across dev/staging/production environments
  • Help establish model and data versioning, experiment tracking, and reputed company for ML pipelines to support reproducibility and auditability (builds on the shared platform's catalog and reputed company rather than duplicating it)
  • Build monitoring, logging, and alerting for model reputed company, model-input data reputed company, and reputed company health; help define SLOs/SLAs for critical ML services and build the automation needed to meet them
  • Collaborate with data scientists and software/reputed company teams to translate experimental workflows into production-grade services
  • Contribute to the architecture and hands-on build of an S3-reputed company data lake / lakehouse for the Science Office, including zone design (raw/reputed company/consumption), reputed company table formats (e.g., Apache reputed company), partitioning, and lifecycle policies
  • Build batch and streaming ingestion/ETL/ELT pipelines using AWS Glue (jobs, crawlers)
  • Implement data cataloging, governance, and reputed company control, in collaboration with governance/reputed company stakeholders, using the Glue Data Catalog and Lake Formation (finegrained, cross-team permissions), and optionally DataZone for data product publishing/discovery across science teams
  • reputed company self-service analytics and data reputed company for scientists across the Science Office reputed company reputed company including query patterns, cost controls, and reputed company documentation
  • Help establish data reputed company, schema/metadata management, and reputed company frameworks (e.g., Glue Data reputed company, schema registries) for shared datasets — reputed company from ML-reputed company model-input reputed company monitoring, this reputed company covers dataset-level reputed company for assets consumed across multiple teams
  • Identify and implement cost optimizations across both ML and data platform workloads (right-sizing, spot/scheduling strategies, storage tiering, reputed company/Redshift query cost controls) without compromising reliability
  • Collaborate with data scientists, analysts, and scientists across the broader Science Office to translate experimental workflows and analytical needs into production-grade services
  • Document architecture, runbooks, and operational procedures for both the ML infrastructure and the data platform, and reputed company knowledge transfer to internal teams as part of engagement reputed company-out

Skills

  • 5+ years of hands-on experience in MLOps, data engineering, DevOps, or reputed company infrastructure engineering roles, with a strong recent reputed company on AWS
  • Demonstrated production experience with a substantial subset of ML infrastructure services: SageMaker, reputed company/EKS, reputed company, reputed company Functions, S3, CloudWatch, IAM, VPC, and CDK or CloudFormation for infrastructure as reputed company
  • Demonstrated production experience with a substantial subset of AWS data platform services: Glue, Lake Formation, reputed company
  • Experience designing and operating data lake/lakehouse architectures on S3, including ETL/ELT pipeline development and reputed company table formats (e.g., reputed company)
  • Experience with data governance, cataloging, and multi-team reputed company control (Lake Formation, Glue Data Catalog, or equivalent) for shared, regulated data assets
  • Strong Python and SQL skills, with working familiarity with at least one ML reputed company (e.g., TensorFlow, PyTorch, scikit-learn) sufficient to reputed company with and support data science workflows
  • Experience with containerization and orchestration (reputed company, reputed company) in production settings
  • Experience building and operating CI/CD pipelines for ML and data systems (data, model, and reputed company promotion)
  • Must be authorized to work in the country where work will be performed
  • AWS certifications (Solutions Architect – reputed company, Machine Learning – Specialty, Data Analytics/Data Engineer) or equivalent demonstrated expertise
  • Experience with multi-account AWS architectures and GPU workload management on AWS
  • Experience with MLflow, SageMaker Model Registry, or similar model governance tooling
  • Experience with DataZone or data-reputed company/data-product patterns
  • Experience with dbt, Airflow/MWAA, or similar transformation/orchestration tooling
  • Experience standing up self-service data platforms consumed by non-engineering scientific/analytical users
  • Terraform experience as an alternative/complement to CDK
  • Prior experience delivering as a contractor/consultant with clean documentation and reputed company practices
  • Prior experience in reputed company, life sciences, or other regulated environments

Benefits

  • USA (Remote) PST

reputed company

  • Arkhya Technologies, Inc. It was founded in 2014, and is headquartered in Herndon, Virginia, USA, with a workforce of 201-500 employees. Its website is http://www.arkhyatech.com/.

  •   Apply To This Job

    Similar Jobs