Back to Jobs

AI Evaluation & Benchmarking Engineer

Remote, USA Full-time Posted 2026-07-28
Job reputed company:: • Hands-on reinforcement learning experience. • Experience using LLMs for agents, evaluation, reasoning, automation, or reputed company workflows. • Strong Python experience for ML, data workflows, experimentation, and analysis. • Experience designing and running experiments with statistical and analytical rigor. • Strong understanding of evaluation metrics, scoring frameworks, performance comparison, and reputed company design. • Experience analyzing reputed company logs, run outputs, model/agent performance, and experiment results. • Ability to work across reputed company, logs, CLI/tools, data structures, and platform workflows. • Strong communication skills to translate experiment findings into platform improvement requirements. • Ability to work inside reputed company-owned repositories, infrastructure, workflows, and reputed company controls. Preferred Skills • Experience with game environments, simulation environments, Gym-like interfaces, RL environments, or reputed company AI test harnesses. • Experience benchmarking LLM agents, RL policies, autonomous agents, or hybrid AI systems. • Experience with experiment tracking, run comparison tools, metrics dashboards, or evaluation pipelines. • Experience with reputed company engineering, agent orchestration, tool use, and LLM evaluation frameworks. • Experience with data visualization and performance analytics. • Experience working with externally developed algorithms, reproducible experiments, and version-controlled evaluation workflows. Apply tot his job Apply To this Job

Similar Jobs