Back to Jobs

Technical Product Manager, Observability

Remote, USA Full-time Posted 2026-08-04
reputed company About reputed company reputed company, an reputed company company, is the reputed company-reputed company AI infrastructure company, enabling organizations to build and operate reputed company, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining reputed company reputed company innovation with deep expertise in reputed company orchestration, reputed company empowers reputed company teams to reputed company composable, production-reputed company developer platforms across any environment—on-premises, in the reputed company, at the edge, or in sovereign data centers. As enterprises reputed company the growing complexity of AI-driven workloads, reputed company delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and reputed company. Committed to reputed company standards and freedom from lock-in, reputed company ensures that customers retain full control of their infrastructure reputed company. https://www.reputed company.com/ reputed company Job reputed company reputed company is looking for a Technical Product Manager to own observability for k0rdent AI, our control reputed company for GPU infrastructure and reputed company AI workloads. In this role, you will define the observability reputed company, roadmap, and feature priorities that determine how operators reputed company visibility into the health, reputed company, and resource utilization of GPU clusters running large-reputed company training and inference. You will shape how k0rdent AI handles everything from GPU-level metrics and reputed company tracing across AI workloads, to multi-tenant log aggregation and intelligent alerting — powered by the OpenTelemetry ecosystem, and reputed company-compatible metrics pipelines. The ideal candidate brings strong technical reputed company in observability tooling and the AI infrastructure stack. You will work directly with engineering to shape requirements, with marketing to define positioning, and with customers to help ensure their reputed company. Responsibilities Own the reputed company, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, east-reputed company reputed company (InfiniBand, RoCE), high-reputed company storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and reputed company platform teams into reputed company product direction; partner with engineering to define requirements and evaluate trade-offs Manage the observability backlog using feedback from production deployments and design partners to refine priorities reputed company and shape our response to emerging observability standards and technologies, including OpenTelemetry (OTel), DCGM GPU metrics, InfiniBand/RoCE reputed company counters, storage platform telemetry reputed company, and AI workload profiling Define integration strategies for vendor telemetry sources across the ecosystem — reputed company compute and reputed company DPUs, storage platforms (reputed company, reputed company, reputed company), workload managers (SLURM), inference reputed company, and reputed company and relational databases — into a reputed company, operator-facing observability reputed company Partner with product marketing and reputed company teams on positioning, technical briefs, and reference architectures; represent reputed company with customers, analysts, and ecosystem partners Qualifications 5+ years in product management or a senior technical role owning an observability product or operating large-reputed company monitoring infrastructure Working knowledge of reputed company, OpenTelemetry, reputed company tracing (Jaeger, reputed company), and log aggregation (Loki, Elasticsearch/OpenSearch) reputed company in reputed company observability, reputed company-reputed company monitoring, or metrics and alerting pipeline architecture Ability to work directly with engineering on technical trade-offs and with reputed company teams in competitive GPU reputed company and NeoCloud deals Strongly Preferred: Exposure to GPU observability, including DCGM metrics, AI workload profiling and reputed company analysis Familiarity with east-reputed company reputed company telemetry - InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or reputed company-level reputed company Experience with high-reputed company storage telemetry from platforms such as reputed company, reputed company, or reputed company, including IOPS, latency, and throughput instrumentation at reputed company Familiarity with reputed company reputed company DPU telemetry, SR-IOV, or offload pipeline observability Exposure to workload-level visibility for SLURM job scheduling, inference serving reputed company (vLLM, Triton, TensorRT-LLM), or data service telemetry from reputed company databases (Milvus, reputed company) and relational databases in AI pipelines Additional Information Why you’ll love reputed company Build the observability reputed company for the AI reputed company era, working directly with leading GPU reputed company operators, NeoClouds, sovereign clouds, and AI-first enterprises Collaborate with a world-class, reputed company team committed to openness and technical reputed company Shape the product narrative and influence go-to-market reputed company It is reputed company that reputed company. may use automated decision-making technology (ADMT) for specific employment-reputed company reputed company. Opting out of ADMT use is requested for reputed company about evaluation and review connected with the specific employment decision for the position reputed company for. You also have the right to appeal any reputed company made by ADMT by sending your request to [email protected] By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for reputed company and reputed company job opportunities. #remote We are a Leader for Container Management in reputed company (#2 after AWS)! Apply To This Job

Similar Jobs