Back to Jobs

[Remote] Senior Research Scientist, Model Evaluation

Remote, USA Full-time Posted 2025-11-24
Note: The job is a remote job and is open to candidates in USA. Cohere is on a mission to scale intelligence to serve humanity by training and deploying frontier models for AI systems. In this role, you will be responsible for creating next-generation evaluation methods and infrastructure to measure LLM progress, pushing the limits of what models can accomplish and ensuring high data quality. Responsibilities • Create ambitious new evaluation benchmarks that push the limits of what our models can accomplish • Work on highly cross-functional teams to translate model feedback into trustworthy, repeatable evaluations • Conduct research to advance the state-of-the-art in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency • Build scalable and reusable tools for digging into model performance Skills • Create ambitious new evaluation benchmarks that push the limits of what our models can accomplish • Work on highly cross-functional teams to translate model feedback into trustworthy, repeatable evaluations • Conduct research to advance the state-of-the-art in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency • Build scalable and reusable tools for digging into model performance • You enjoy rapidly building prototypes that demonstrate the boundaries of what LLMs are capable of, and you have developed resources to measure those capabilities • You have spent dozens of hours reviewing complex data and LLM outputs to ensure high data quality • You are obsessive about rigorously measuring AI capabilities, and also about making sure your measurements actually align with the capabilities you care about • You have strong software engineering skills Benefits • An open and inclusive culture and work environment • Work closely with a team on the cutting edge of AI research • Weekly lunch stipend, in-office lunches & snacks • Full health and dental benefits, including a separate budget to take care of your mental health • 100% Parental Leave top-up for up to 6 months • Personal enrichment benefits towards arts and culture, fitness and well-being, quality time, and workspace improvement • Remote-flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co-working stipend • 6 weeks of vacation (30 working days!) Company Overview • Cohere is an enterprise AI firm developing secure and private AI technology to address real-world business challenges. It was founded in 2019, and is headquartered in Toronto, Ontario, CAN, with a workforce of 201-500 employees. Its website is https://cohere.com. Company H1B Sponsorship • Cohere has a track record of offering H1B sponsorships, with 11 in 2025, 14 in 2024, 13 in 2023, 5 in 2022, 2 in 2021. Please note that this does not guarantee sponsorship for this specific role. Apply tot his job Apply To this Job

Similar Jobs

[Remote] E-commerce Product Manager (Contract)

Remote, USA Full-time

SQL Developer

Remote, USA Full-time

AI Engineer Intern

Remote, USA Full-time

AI-Based Cybersecurity Research Intern

Remote, USA Full-time

Data Science and Analytics Senior Manager (Virtual)

Remote, USA Full-time

Business Analyst

Remote, USA Full-time

Senior Manager, CRM Systems Administration

Remote, USA Full-time

Credit Adjudicator

Remote, USA Full-time

[Remote] 5G RAN Systems Engineer

Remote, USA Full-time

Vendor Management Specialist – Momentum Manufacturing Group – North LLC – Georgetown, MA

Remote, USA Full-time

[Remote/WFM] Data Entry Jobs Flexible Schedule, Work Remotely

Remote, USA Full-time

Accountant - Chicago, IL - Full-Time

Remote, USA Full-time

LVN II-Kern San Dimas MOB-Virtual MC Appointment Services-Part Time

Remote, USA Full-time

[Remote/WFM] Data Entry Clerk Work From Home - Part Time Focus

Remote, USA Full-time

Join Today: Urgently Need Math Instructor / Tutor in Lake

Remote, USA Full-time

SNF Placement Case Manager (LVN, RN – Part Time) CA

Remote, USA Full-time

**Experienced Medical Assistant – Remote Customer Service Representative at blithequark**

Remote, USA Full-time

[Remote/WFM] Data Entry Operator / Entry Level (Remote)

Remote, USA Full-time

Experienced Customer Service Representative – After Hours Support Specialist for Healthcare and Transportation Services – Remote Opportunity in Virginia

Remote, USA Full-time

Immediately Require School Year Tutor in Garfield, NJ

Remote, USA Full-time