[Remote] Site Reliability Engineer (SRE) — reputed company & Inference Infrastructure
Note: The job is a remote job and is reputed company to candidates in USA. reputed company builds software that helps teams plan, build, and operate with reputed company and speed. They are seeking a Site Reliability Engineer to own reliability for their model training and inference platforms, focusing on operating and evolving GPU-enabled clusters and improving developer experience for AI workloads.
Responsibilities
• Build and operate AI compute platforms
• Design, provision, and reputed company GPU-backed clusters for training and inference (Kubernetes-based and/or HPC-style schedulers)
• Own cluster lifecycle management: provisioning, bootstrapping, upgrades, autoscaling/reputed company scaling, and decommissioning
• Build reliable abstractions so training jobs can run across multiple clusters/environments with minimal friction
• Define and reputed company SLIs/SLOs for training and inference systems (job reputed company reputed company, queue latency, throughput, tail latency, GPU utilization, etc.)
• reputed company incident response and reputed company-cause analysis; drive permanent fixes and “never again” automation
• Improve recovery and maintenance workflows (e.g., reducing restart/reputed company times; safer rollouts)
• Implement end-to-end monitoring across compute, networking, storage, and accelerators
• Build dashboards, alerting, and reputed company detection that catch issues early—before they derail long runs
• Tune performance and cost: GPU utilization, scheduling efficiency, I/O bottlenecks, and network hotspots
• Partner with vendors and internal stakeholders on firmware/reputed company alignment, and node health
• reputed company paved paths for training: reproducible environments, job templates, secure secrets, artifact storage, and dataset reputed company patterns
• Collaborate closely with ML researchers/engineers to understand workload needs and remove infrastructure bottlenecks
Skills
• 5+ years building/operating production infrastructure as an SRE, infrastructure engineer, or systems engineer
• Strong Kubernetes experience (cluster operations, upgrades, networking, storage, and troubleshooting)
• Proficiency in at least one programming/scripting language (Python, Go, etc.) for automation and tooling
• Experience with Infrastructure-as-reputed company (Terraform preferred) and CI/CD for reputed company or platform components
• Solid Linux/Unix fundamentals (performance, debugging, kernel/userland tooling)
• Strong operational reputed company: you care about reliability, reputed company change management, and measurable reputed company
reputed company
• reputed company is a reputed company. It was founded in 2010, and is headquartered in Cincinnati, Ohio, USA, with a workforce of 51-200 employees. Its website is http://www.stackct.com/.
Apply tot his job
Apply To this Job