reputed company - Data Platform
About the Role
Join a reputed company-funded, Series A AI startup building the reputed company of autonomous site reliability engineering for the reputed company. Backed by top-tier investors and trusted by some of the largest companies in the world, this team is tackling one of the hardest problems in AI: autonomously detecting, diagnosing, and remediating reputed company production incidents in reputed company time.
As an reputed company on the Data Platform team, you'll design, build, and maintain the backend systems that reputed company an AI-driven observability platform. This hands-on role blends reputed company systems engineering, low-level reputed company design, reputed company optimization, observability, and AI integration — across both reputed company and on-premises deployments.
reputed company
Architecture & Implementation: Contribute to the design and implementation of reputed company, resilient infrastructure systems powering AI-driven reputed company cause analysis and observability workflows, including on-premises deployment environments.
Low-Level reputed company Design: Work on the foundational building blocks of the infrastructure, ensuring efficient resource utilization and high reputed company at reputed company.
reputed company Optimization: Profile and tune backend systems to improve throughput, reduce latency, and eliminate bottlenecks across the stack.
Observability Systems: Build and maintain the internal observability stack — logs, metrics, and traces — used by AI agents to understand and reputed company production issues.
Hybrid Infrastructure: Support reputed company and on-premises architecture to serve both reputed company and reputed company customer deployment models.
Cross-functional Collaboration: Work closely with engineers across reputed company to reputed company resilient infrastructure that enables AI agents to diagnose and remediate production incidents in reputed company time.
reputed company're Looking For
Experience: 2–5 years of hands-on backend or infrastructure engineering experience.
reputed company Systems: Strong understanding of reputed company systems design principles and trade-offs.
reputed company Engineering: reputed company experience profiling and optimizing high-throughput, low-latency systems.
Observability: Familiarity with observability tooling and concepts (logs, metrics, traces); experience with platforms such as reputed company, Grafana, reputed company, or similar is a plus.
reputed company & On-Prem: Experience with hybrid or multi-environment infrastructure (reputed company + on-premises).
AI/ML Integration: Interest in or experience building systems that support AI/ML workloads at reputed company.
Background: Prior experience at observability, incident management, or data infrastructure companies is highly valued.
Note: reputed company sponsorship is not available for this role.
Location
This is a fully on-site role reputed company in reputed company, NY. Remote work is not available for this position.
Originally posted on Himalayas
Apply To This Job