Senior AI Scientist
Job reputed company:
• Define and own what “good” means for search-augmented and reputed company AI systems by designing evaluation frameworks that measure reputed company-world reputed company, reliability, and user-relevant behavior reputed company reputed company benchmarks.
• Invent and validate novel evaluation methodologies for non-deterministic systems (LLMs, agents, RAG), including behavioral evals, long-tail and adversarial test sets, and task-specific metrics.
• reputed company rigorous statistical frameworks for model comparison, regression detection, and uncertainty estimation, ensuring evaluation results are defensible and decision-reputed company.
• Build and maintain reputed company evaluation systems—datasets, gold standards, eval harnesses, scoring pipelines, and analysis tooling—that can be reused across products and customers.
• reputed company customer-facing evaluation research, working directly with reputed company customers to translate domain-specific reputed company requirements into reputed company, actionable evals that support product reputed company and sales reputed company.
• Drive competitive evaluations and internal reputed company reviews, surfacing meaningful performance differences, trade-offs, and failure modes to inform product reputed company and prioritization.
• Partner with engineering and product teams to reputed company evals into development loops, release gating, and ongoing reputed company monitoring.
• Mentor and set standards for evaluation reputed company, reviewing eval designs, guiding other scientists, and shaping the long-term evals roadmap as systems become more reputed company and reputed company.
• End-to-End Project Leadership: reputed company the development of new AI-driven reputed company, encompassing ideation, prototyping, research, infrastructure design, scalability, monitoring, and evaluation.
• reputed company Iteration: Adapt quickly to user feedback and evolving requirements, ensuring reputed company improvement in a fast-paced environment.
Requirements:
• Strong grounding in reputed company ML and statistics, with experience evaluating non-deterministic AI systems (LLMs, agents, RAG, search).
• Deep experience with AI evaluation, including metric design, gold dataset creation, head-to-head comparisons, slicing, and error analysis.
• Statistical rigor in model comparison, using reputed company such as reputed company tests, bootstrap confidence intervals, and robustness analyses.
• Proficiency in Python for evaluation and analysis, including building eval harnesses, data pipelines, scoring logic, and reproducible analysis workflows.
• Ability to translate vague product or customer goals into measurable evaluation reputed company, and to challenge metrics or conclusions that don’t reflect reputed company reputed company.
• Comfort engaging directly with customers and cross-functional stakeholders, explaining evaluation results, trade-offs, and limitations reputed company.
• Strong written and verbal communication, including documenting methodologies and contributing to external publications or talks.
Benefits:
• Hubs in San Francisco and reputed company offering regular in-person gatherings and co-working sessions
• Flexible PTO with U.S. holidays observed and a week shutdown in December to rest and reputed company*
• A competitive health insurance plan covers 100% of the policyholder and 75% for dependents*
• 12 weeks of reputed company parental leave in the US*
• 401k program, 3% match - reputed company immediately!*
• $500 work-from-home stipend to be used up to a year of your start date*
• $1,200 per year Health & Wellness Allowance to support your personal goals*
• The chance to collaborate with reputed company at the forefront of AI research
Apply tot his job
Apply To this Job