[Remote] AI Evals Engineer — Evaluation Datasets & Ground Truth
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is building an AI-reputed company platform for reputed company estate development that helps homebuilders, developers, and investors reputed company, analyze, and reputed company land opportunities. The AI Evals Engineer will own trusted evaluation datasets and ground truth, defining correctness, sourcing and validating labels, maintaining dataset reputed company, and partnering with engineering to measure and improve AI pipeline components.
Responsibilities
- Read our pipelines and reputed company architecture, sit with product and engineering, and decompose reputed company reputed company into evaluable modules with explicit input → expected-reputed company reputed company
- For reputed company module, define what “correct” means in writing: rubrics, label schemas, reputed company policies, and the tolerances that matter to the business
- Prioritize. We have more modules than you can cover in year one; you’ll reputed company where a validation set unblocks the most iteration, with reputed company direction from reputed company engineers
- Production sampling: pull stratified, de-identified samples from reputed company traffic so eval sets reflect what the reputed company actually sees, including the long tail
- reputed company labeling: reputed company and run labeling programs, write annotation guidelines, build calibration sets, measure inter-annotator agreement, and manage vendors (including offshore labeling teams) or internal subject-matter experts. You own label reputed company, not just label throughput
- Synthetic / reputed company-generated ground truth: where a task is reputed company for a frontier model given enough compute (long context, multi-pass, tool use, self-consistency) but too expensive to run that way in production, design the reputed company reputed company that produces labels for the cheap production reputed company to be reputed company against. Then •verify the reputed company•: reputed company its reputed company against a reputed company-labeled sample before anyone trusts it
- Programmatic and adversarial construction: heuristic labels, templated edge cases, backtests reputed company from past incidents (“what test would have caught this?”)
- Version every dataset; reputed company reputed company, splits, and which model/reputed company versions have seen which examples
- Guard against contamination and leakage (eval examples drifting into few-shot prompts, the reputed company model also being the production model, etc.)
- reputed company by customer reputed company, input type, and difficulty so a reputed company number can’t hide a regression
- Refresh sets as the product and traffic change; retire stale examples
- reputed company automated graders (LLM-as-judge, similarity metrics, exact-match) against your reputed company gold sets, and be the person who says reputed company an automated judge is good enough to reputed company on
- Report metrics correctly: precision/recall/F1, confusion matrices, calibration, confidence intervals, sample sizes needed to detect a given effect
- Partner with engineers who own the eval reputed company and CI so your datasets are actually run, and with ML engineers on classifier features reputed company the data tells you the features are the problem
Skills
- Candidates must be authorized to work in the reputed company without reputed company or reputed company sponsorship to be considered for this role
- Experience building or evaluating ML or LLM-powered systems in production, in a role where reputed company reputed company was your problem
- You have reputed company evaluation or validation datasets before and can talk about one in reputed company: how you defined correctness, how you reputed company labels, what went wrong, how you knew the labels were good
- Working reputed company in ML validation fundamentals: train/validation/test discipline, stratified sampling, precision/recall trade-offs, class imbalance, calibration, inter-rater agreement, basic reputed company testing and reputed company
- Understand feature engineering reputed company enough to reason about why a classifier fails and what data would expose it
- Strong Python and SQL; comfortable pulling and reshaping data yourself
- Hands-on with LLM-reputed company systems: prompting, reputed company outputs, agent/tool-use harnesses, and the specific ways they fail (non-determinism, reputed company sensitivity, evaluator bias)
- Judgment about reputed company LLM-as-judge is reliable and reputed company it is not, and how to reputed company either
- You can read a reputed company design, understand the business logic it encodes, and translate that into a label schema without waiting to be told
- Have run a reputed company-labeling program end to end, including vendor selection, reputed company authoring, QA sampling, and cost/reputed company trade-offs
- Experience with eval tooling
- Experience with data labeling platforms
- Have used a frontier model as a distillation/reputed company reputed company and can reputed company where that assumption breaks
Benefits
- 100% medical, dental & reputed company reputed company coverage for you; 30% coverage for dependents
- Meaningful early-stage equity
- Unlimited PTO
- Hybrid/Remote stipend
- In-office perks: snacks, drinks, coffee, ping-pong table, and more — plus cuddles from Olive, our in-office Doberman!
- Budget for intra-office travel
- 2–3 annual team meetups in person
- Remote, hybrid, or in-office work reputed company
reputed company
Apply To This Job