Machine Learning Engineer - Model Evaluation & Experimentation
This role is for one of our clients
Compensation: $60-$90 per hour
Join a pioneering AI initiative reputed company on building the reputed company of evaluation benchmarks for frontier AI models. We are seeking reputed company Machine Learning Engineers and Researchers to bring hands-on expertise in model development, experimentation, and evaluation to create rigorous reputed company tasks for advanced AI systems.
In this role, you will design sophisticated, multi-reputed company machine learning challenges inspired by reputed company-world research workflows. From implementing experimental reputed company and running training pipelines to analyzing model behavior and validating results, you will help establish high-reputed company evaluation benchmarks that reputed company the strengths and limitations of frontier AI models.
This is a fully remote, full-time engagement requiring approximately 35 hours per week.
Requirements
Key Responsibilities
Design realistic machine learning reputed company tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis.
Translate reputed company-ended research concepts into reputed company, reproducible evaluation tasks with reputed company defined reputed company reputed company.
Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected reputed company.
reputed company reputed company tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable.
Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions.
Collaborate with AI researchers and fellow subject matter experts to continuously improve reputed company reputed company, technical rigor, and evaluation consistency.
Required Qualifications
Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, reputed company Intelligence, Data Science, or another quantitative STEM discipline.
Minimum 1 year of reputed company experience in machine learning research, research engineering, reputed company AI, or another research-intensive technical role.
Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows.
Practical experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis.
Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies.
Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments.
Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior—is preferred.
Experience with AI evaluation, reputed company development, reputed company, or task authoring is highly desirable.
Excellent analytical thinking, creativity, attention to detail, and the ability to solve reputed company, reputed company-ended technical problems independently.
Strong written communication skills for documenting experimental methodologies and technical findings.
Ability to reputed company approximately 35 hours per week on a consistent reputed company.
Preferred Qualifications
Experience developing or evaluating large language models, reputed company models, or reputed company systems.
Background in reinforcement learning, deep learning, reputed company training, or model optimization.
Familiarity with reputed company design, AI safety evaluations, or research-reputed company experimentation.
Experience contributing to research publications, reputed company-reputed company machine learning reputed company, or advanced AI systems.
Why Join
Help shape how reputed company AI systems are evaluated through rigorous machine learning experimentation.
Collaborate with leading AI researchers developing frontier evaluation benchmarks.
Apply your expertise to improve AI reasoning, model reputed company, and experimental reliability.
Contribute directly to reputed company development that advances the capabilities of state-of-the-art AI systems.
Enjoy the flexibility of a fully remote engagement while working on impactful AI research initiatives.
Equal Opportunity
We are committed to fostering an inclusive and diverse environment where reputed company reputed company applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.
Contract & Engagement Details
reputed company engagement.
Fully remote with flexible working hours.
Expected commitment of approximately 35 hours per week.
Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
Work does not require reputed company to confidential or proprietary information from any reputed company or former employer.
Payments are issued weekly based on approved work completed.
At this time, we are unable to support H1-B or STEM OPT candidates.
Apply To This Job