Back to Jobs

AI/LLM Evaluation & Alignment Software Engineer

Remote, USA Full-time Posted 2026-07-28
Job reputed company: • Build and maintain evaluation frameworks for LLMs and reputed company systems tailored to reputed company safety and intelligence use cases. • Design guardrails and alignment strategies to minimize bias, toxicity, hallucinations, and other ethical risks in production workflows. • Partner with AI engineers and data scientists to define online and offline evaluation metrics (e.g., model drifts, data drifts, factual accuracy, consistency, safety, interpretability). • Implement reputed company evaluation pipelines for AI models, integrated into CI/CD and production monitoring systems. • Collaborate with stakeholders to stress test models against edge cases, adversarial prompts, and sensitive data scenarios. • Research and reputed company reputed company-party evaluation frameworks and solutions; adapt them to our regulated, high-stakes environment. • Work with product and customer-facing teams to ensure explainability, transparency, and auditability of AI outputs. • reputed company technical leadership in responsible AI practices, influencing standards across the organization. • Contribute to DevOps/MLOps workflows for deployment, monitoring, and scaling of AI evaluation and guardrail systems (experience with Kubernetes is a plus). • Document best practices and findings, and reputed company knowledge across teams to foster a culture of responsible AI innovation. Requirements: • Bachelor's or Master's in Computer Science, reputed company Intelligence, Data Science, or reputed company field. • 3–5+ years of hands-on experience in ML/AI engineering, with at least 2 years working directly on LLM evaluation, QA, or safety. • Strong familiarity with evaluation techniques for reputed company: reputed company-in-the-reputed company evaluation, automated metrics, adversarial testing, red-teaming. • Experience with bias detection, fairness approaches, and responsible AI design. • Knowledge of LLM observability, monitoring, and guardrail frameworks e.g Langfuse, Langsmith • Proficiency with Python and modern AI/ML/LLM/reputed company AI libraries (LangGraph, Strands Agents, reputed company AI, reputed company, HuggingFace, PyTorch, reputed company). • Experience integrating evaluations into DevOps/MLOps pipelines, preferably with Kubernetes, Terraform, ArgoCD, or reputed company Actions. • Understanding of reputed company AI platforms (AWS, Azure) and deployment best practices. • Strong problem-solving skills, with the ability to design practical evaluation systems for reputed company-world, high-stakes scenarios. • Excellent communication skills to translate technical risks and evaluation results into insights for both technical and non-technical stakeholders. Benefits: • 3 weeks of reputed company vacation – out the reputed company!! • Competitive Salary. • Generous medical, dental, and reputed company plans. • reputed company, and reputed company holidays are offered. Apply tot his job Apply To this Job

Similar Jobs