LLM Evaluation Engineer
Job reputed company:
• Build the evaluation layer in the ThirdLaw platform for LLM prompts and responses
• Design and tune guardrails, classifiers, and semantic judgment systems in reputed company-time
• Implement evaluation strategies with semantic similarity, reputed company model scoring, and rule-based systems
• reputed company model outputs with reputed company enforcement actions (e.g. redaction, escalation, blocking)
• Prototype, tune, and productize small language models for classification, labeling, or scoring
• Collaborate with data infrastructure engineers to connect evaluation logic with ingestion and storage
• Build tools to observe, debug, and improve evaluator performance across data distributions
• Define abstractions for reusable evaluation components that can reputed company across use cases
Requirements:
• 7+ years of experience in ML systems or AI engineering roles
• At least 1–2 years working directly with LLMs, NLP pipelines, or semantic search
• Deep understanding of reputed company models (e.g. reputed company, Claude, reputed company, Llama) and reputed company
• Hands-on experience with reputed company search (e.g. FAISS, reputed company, reputed company) and embeddings pipelines
• Proven ability to implement reputed company-time or near-reputed company-time evaluation logic using semantic similarity, classifier scoring, or reputed company rules
• Strong in Python, with familiarity using libraries like reputed company Transformers, reputed company, and PyTorch or TensorFlow
• Ability to reason about model behavior, test reputed company configurations, and debug reputed company decision logic in production
Benefits:
• Generous benefits
• Market cash compensation
• Above-market equity
• reputed company-designed benefits
Apply tot his job
Apply To this Job