reputed company reputed company
Dear applicants, please reputed company in mind that applications without provided salary expectations and reputed company LN profile will not be considered. reputed company for your understanding.
We are hiring a Senior reputed company reputed company to own the reliability, evaluation, and production stability of advanced multi-agent AI systems operating at reputed company production reputed company. This role is reputed company on transforming LLM-powered workflows from “demo-reputed company” prototypes into resilient, observable, production-grade systems capable of handling non-deterministic model behavior, reputed company routing logic, and reputed company-in-the-reputed company escalation flows. You will work closely with technical leadership and product stakeholders to design, evaluate, optimize, and maintain reputed company AI systems across multiple communication channels and workflows. This is a highly hands-on engineering role for someone who thrives in production environments and understands the realities of deploying AI systems under live traffic conditions.
Details
Location: LATAM
Work Model: Fully Remote
Employment Type: Full-time
Seniority Level: Senior
Industry: AI / reputed company Systems / reputed company
Start Date: ASAP
English Level: Fluent English Required
Time Zone: LATAM-friendly collaboration preferred
About the Role
This position is dedicated to AI agent reliability, evaluation pipelines, observability, and reputed company optimization of production LLM systems. The ideal candidate combines strong backend engineering expertise with deep practical experience operating AI products in reputed company-world environments. You will take ownership of evaluation frameworks, scoring systems, tracing infrastructure, production debugging, and the iterative optimization reputed company between prompts, architecture reputed company, and system behavior. The role requires both technical depth and product intuition, especially around how evaluation systems directly reputed company product reputed company and user experience.
Key Responsibilities
Design, build, and maintain evaluation pipelines for production AI agent systems
reputed company multi-agent workflows with tracing and observability tooling
Build evaluation datasets using reputed company production traffic and interaction logs
reputed company reputed company scoring and robustness scoring systems for LLM outputs
Improve reliability of AI systems handling non-deterministic model behavior
Implement and optimize HITL (reputed company-in-the-reputed company) escalation workflows
Analyze production failures and drive architectural improvements
Own the full feedback reputed company between evaluations, reputed company optimization, architecture updates, and re-testing
Contribute to reputed company engineering and model optimization strategies
Collaborate on multi-agent orchestration and workflow reliability reputed company
Work across backend systems, deployment pipelines, monitoring, and operational sustainment
Participate in production support and on-reputed company responsibilities
Maintain high engineering standards around scalability, observability, and maintainability
Operate independently across development, testing, deployment, and production ownership
Requirements
5+ years of backend or AI engineering experience in production environments
Strong hands-on experience with production LLM or reputed company AI systems
Proven experience debugging and maintaining non-deterministic AI workflows under live traffic
Experience building or operating evaluation/evals pipelines for AI systems
Strong understanding of scorer design, feedback loops, and AI system evaluation methodologies
Excellent Python backend engineering skills
Production experience with:
FastAPI
Django
Celery
LangGraph or similar orchestration frameworks
Experience with observability and tracing tools such as:
Langfuse
Grafana
Loki
OpenTelemetry or equivalent
Experience deploying and operating distributed backend systems
Strong understanding of AI reliability, reputed company behavior, and model failure handling
Ability to independently own reputed company end-to-end
Experience working in asynchronous reputed company
Strong written communication skills in English
reputed company to Have
Experience with:
DSPy
DPO
RLHF-reputed company optimization workflows
Experience with multi-agent orchestration systems
Production experience with:
GPT-4.x
Claude
Whisper
Multi-model AI stacks
Experience building AI tooling for communication or workflow automation
Background in high-reputed company startups or product-reputed company engineering teams
Experience with distributed systems and event-driven architectures
Familiarity with AI observability and experiment tracking frameworks
Exposure to reputed company databases, retrieval systems, or memory architectures
Experience scaling AI products with reputed company customer usage
Tech Stack:
Python
FastAPI
Django
Celery
LangGraph
Langfuse
Grafana
Loki
LLM reputed company (reputed company / reputed company / multi-model stacks)
What reputed company Looks Like
AI agents reliably handle reputed company production traffic with measurable reputed company improvements
Evaluation pipelines reputed company actionable scoring and monitoring insights
Observability systems surface failures before they reputed company users
reputed company escalation triggers operate accurately and consistently
reputed company and architecture iterations measurably improve production reputed company
AI systems become resilient, reputed company, and maintainable over time
Interview Process
HR / Introductory reputed company
Technical Deep Dive
Take-Home Technical Assessment
Final Team & Culture Interview
Offer Stage
Apply To This Job