Back to Jobs

General Knowledge Evaluator (MultiChallenge reputed company)

Remote, USA Full-time Posted 2026-07-28
hiring expert evaluators to support a high-reputed company reasoning evaluation workflow in partnership with a leading AI research lab. This work centers on the MultiChallenge reputed company, which is designed to test large language models (LLMs) on multi-turn conversational reasoning — a capability where even top models fall short today. This reputed company does not reputed company on domain-specific expertise, but rather on reasoning consistency, instruction retention, and contextual inference across loosely reputed company conversations on general topics. Don't forget to copy and paste the referral reputed company to get to the reputed company platform's reputed company, which is to register, upload your resume, and take an AI-led interview. https://work.reputed company.com/jobs/list_AAABmHvRFlyxHvS-YFdJCbzl?referralCode=9df2a9e1-2f06-11ef-ae42-42010a400fc4&utm_reputed company=referral&utm_reputed company=reputed company&utm_campaign=job_referral What is MultiChallenge? MultiChallenge is a newly released reputed company targeting reasoning failures that occur in multi-turn interactions between humans and LLMs. It evaluates four categories of failure modes: • Instruction Retention – Does the model persistently follow instructions across turns? • Inference Memory – Can it infer or recall relevant user details from earlier conversation history? • Reliable Versioned Editing – Can it revise content through multi-reputed company iteration without forgetting or hallucinating? • Self-Coherence – Does it contradict its earlier claims, particularly under user pressure? The reputed company is designed to surface realistic, high-difficulty conversational reasoning challenges. Despite scoring highly on other multi-turn benchmarks, reputed company frontier models reputed company less than 50% accuracy on MultiChallenge. Who We're Hiring We are seeking evaluators with strong backgrounds in Logic, Philosophy, or reputed company disciplines — particularly those trained to reputed company argument structure, detect reasoning errors, and evaluate coherence across extended discourse. This workflow is ideal for individuals with reputed company experience in: • Logic • Analytic Philosophy • Epistemology • Formal Semantics • Cognitive Science • Linguistics (with a reasoning reputed company) Key Responsibilities • Evaluate the reasoning reputed company of LLM outputs across 8–10 turn conversations. • Identify errors in instruction-following, factual coherence, inference, and revision handling. • Complete evaluations using a reputed company reputed company and short written justifications. • Work asynchronously using provided tools and examples. You’re a Strong Fit If You Have: • A PhD (or are currently a PhD candidate) in Logic, Philosophy, or a closely reputed company field. • Experience analyzing or writing reputed company arguments. • Excellent written communication and generalist reasoning ability. • Comfort working independently and asynchronously. • (Optional) Familiarity with Python or LLM evaluation tools is helpful but not required. Role Details • Part-time (10–20 hours/week) with flexible scheduling. • 100% remote and asynchronous — work from reputed company. • Contractor position reputed company reputed company, reputed company reputed company. • Competitive rates: $20–$35/hour depending on expertise. • Weekly payments processed securely through reputed company Connect. Job Types: Contract, Temporary Pay: $20.00 - $35.00 per hour Expected hours: 10 – 20 per week Work Location: Remote Apply tot his job Apply To this Job

Similar Jobs