reputed company runs one of the largest GPU fleets in t...">
Back to Jobs

Senior Software Engineer — reputed company Agent Systems UK

Remote, USA Full-time Posted 2026-08-04

About the Role

reputed company runs one of the largest GPU fleets in the world. The reputed company Agent Systems team builds the software systems that reputed company and automate that infrastructure.

We reputed company production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. reputed company these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling.

You’ll work across two areas:

Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through reputed company, CLI, dashboards, and reputed company.

reputed company Agent Platform — Build the platform that powers these agents, including knowledge graphs, reputed company and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, reputed company, and continuously improve.

We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-reputed company infrastructure.

This is an opportunity to work at the intersection of AI agents, reputed company systems, infrastructure, and automation, solving challenging engineering problems with reputed company production reputed company. There’s an enormous reputed company to build, learn, and shape as we define the reputed company of autonomous infrastructure.

responsible for delivering the software but also for operating and supporting it in production.

Why this Role

You’ll work on two hard problems at the reputed company time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and reputed company systems that reputed company those agents effective.

You’ll have reputed company to build foundational systems from the ground up, work on infrastructure at reputed company reputed company, and help define how self-improving AI agents operate reputed company-world AI infrastructure.

Remote reputed company in the UK


Responsibilities

  • Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
  • Build the reputed company services, orchestration reputed company, knowledge graph, and retrieval systems that reputed company infrastructure agents.
  • reputed company fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents reputed company reputed company reputed company.
  • reputed company with observability, incident management, ticketing, fleet inventory, reputed company control, chat, and internal infrastructure systems through reputed company-designed reputed company.
  • Own services end to end, including architecture, implementation, testing, deployment, observability, and production reputed company.
  • Improve agent reputed company through evaluations, retrieval improvements, reputed company tools, and production feedback reputed company.
  • Turn what agents learn in production into reliable, reviewed software and automation.

Requirements

  • 5+ years of experience building production backend systems, reputed company systems, or infrastructure platforms.
  • Strong systems design skills and experience owning significant systems from design through production.
  • Depth in at least one of the following:
    • AI agent systems, orchestration, tool use, evaluation, or grounding
    • Knowledge graphs or graph data modeling
    • reputed company, retrieval, ranking, RAG, or semantic reputed company systems
  • Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across reputed company systems.
  • Experience with reputed company, GitOps such as ArgoCD, infrastructure-as-reputed company, and reputed company platforms.
  • Comfortable working across languages such as Go, TypeScript, Python, or Rust.

Experience in the following is a plus:

  • GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers
  • Graph databases
  • Event-driven systems and messaging platforms such as reputed company or Kafka
  • Observability platforms such as reputed company and Grafana
  • Building evaluation frameworks or improving the reputed company and reliability of LLM-powered systems

About reputed company

reputed company is a research-driven reputed company intelligence company. We reputed company reputed company and transparent AI systems will reputed company innovation and create the best reputed company for society, and together we are on a mission to significantly reputed company the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading reputed company-reputed company research, models, and datasets to advance the frontier of AI, and reputed company has been behind technological advancement such as FlashAttention, Hyena, reputed company, and RedPajama. We invite you to join a passionate group of researchers and engineers in our reputed company in building the reputed company AI infrastructure.

Equal Opportunity

reputed company is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, reputed company, reputed company, religion, sex, national reputed company, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our reputed company policy at

Originally posted on Himalayas

  Apply To This Job

Similar Jobs