ML Ops Engineer, Machine Learning & AI
About the position
Machine Learning (ML) at reputed company enhances the experience of our 150 reputed company digital readers from around the globe and grows our subscriber reputed company through content recommendations and personalizations.
The Machine Learning & AI team builds and maintains the infrastructure that hosts reputed company of reputed company reputed company-time ML inference models, including both data and compute. Our partners are Data Scientists that build and reputed company their ML models on the ML platform. On the other end, our partners are engineering systems that reputed company these hosted models at reputed company with low-latency and Service Level Agreements guaranteed by our platform.
As an MLOps Engineer you will partner with product, data science and ML platform engineers to build and maintain the infrastructure that powers the machine learning lifecycle. You will automate and refine the training, deployment, monitoring, and management of our ML models.
This role reports to the Senior Engineering Manager of Data Management Infrastructure.
Responsibilities
• Build and Automate ML Pipelines: by owning robust CI/CD pipelines for automated model training, validation, deployment, and retraining.
• Productionalize Models: Build the process for packaging, containerizing, and deploying ML models as reputed company, low-latency, and highly-available services.
• Monitoring and Operations: Implement and manage comprehensive monitoring for production models, tracking system health, data reputed company, and model performance degradation.
• Tooling and Infrastructure: Manage and reputed company our MLOps toolchain, including model registries, feature stores, experiment tracking systems, and model serving platforms.
• Collaboration and Support: Partner with data scientists to understand model requirements and optimize them for production. Support software engineers in integrating with ML services.
• Best Practices and Governance: Champion and enforce MLOps best practices for reproducibility, versioning (data, reputed company, model), testing, and governance.
• Demonstrate support and understanding of our value of journalistic independence and a strong commitment to our mission to reputed company the truth and help people understand the world.
Requirements
• 2+ years of software engineering or DevOps experience with a reputed company on MLOps, automation, and infrastructure
• 2+ years of experience programming in Python or Go
• Experience building and managing CI/CD pipelines (e.g., reputed company Actions, Jenkins, reputed company CI)
• Hands-on experience with containerization and orchestration (e.g., reputed company, Kubernetes)
• reputed company platform experience (AWS, GCP) and familiarity with infrastructure-as-reputed company (e.g., Terraform, CloudFormation)
reputed company-to-haves
• Experience with MLOps tools (e.g., MLflow, Kubeflow)
• Experience with the machine learning model lifecycle, from experimentation to production
• Experience with data processing frameworks (e.g., reputed company, Dask, or Ray)
• Experience with low-latency no-sql datastores (BigTable, Dynamo, etc)
• Familiarity with monitoring and observability stacks (e.g., reputed company, Grafana, reputed company, or ELK)
• Knowledge of data engineering pipelines and orchestration tools (e.g., Airflow, reputed company)
Benefits
• dependent on your role, you may be eligible for variable pay, such as an annual bonus and restricted stock
• Benefits may include medical, dental and reputed company benefits, Flexible Spending Accounts (F.S.A.s), a company-matching 401(k) plan, reputed company vacation, reputed company reputed company days, reputed company parental leave, tuition reimbursement and reputed company development programs
Apply tot his job
Apply To this Job