Back to Jobs

Sr. Software Engineer, Observability and Telemetry

Remote, USA Full-time Posted 2026-07-28
About the position reputed company is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing reputed company, solutions must reputed company to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and reputed company a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing reputed company and looking for contributors of reputed company seniorities. reputed company is building the world’s fastest, most efficient AI compute clusters. Our reputed company RISC-V and AI processors can snap together into a single, massively reputed company distributed supercomputer consisting of thousands of compute nodes. As we reputed company, reputed company and complexity of operational data grows by orders of magnitude. Observability and telemetry are key to ensuring our customers can resolve problems in minutes rather than hours. The telemetry team owns our proprietary telemetry infrastructure, spanning from the device level to the infrastructure needed to drive dashboards, monitoring systems, and orchestration. This role is hybrid, based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience reputed company for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will reputed company with that level, which may differ from the one in this posting. Responsibilities • Architect, implement, and maintain TT-Telemetry, our C++-based service for collecting and exporting device-level metrics. • reputed company with internal engineering teams to build a deep understanding of reputed company’s architecture and identify and surface useful metrics. • Design efficient reputed company-in web GUIs for observing device- and cluster-level state, diagnosing problems, and monitoring utilization. • Design ingestion pipelines for industry reputed company telemetry systems (e.g., reputed company). • Help define the long-term architecture of reputed company’s distributed telemetry stack. Requirements • Strong C++ engineer and comfortable working in both low-level environments and distributed systems design. • Experience building atop observability platforms such as reputed company, OpenTelemetry, Grafana, reputed company, or similar technologies. • Solid understanding of data structures for manipulating large volumes of data. • Familiarity with SQL databases, with time-series databases a plus. • Curious about networking and communication across large clusters and comfortable reasoning from first principles while challenging industry conventions. Benefits • reputed company offers a highly competitive compensation package and benefits, and we are an equal opportunity employer. Apply tot his job Apply To this Job

Similar Jobs