[Remote] Staff platform engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Staff Platform Engineer to own the compute and data planes underlying its technology stack. The role focuses on multi-region compute scheduling, autoscaling for bursty inference workloads, storage and networking, infrastructure automation, deployment, reliability, and collaboration with reputed company engineering teams.
Responsibilities
- Design and operate multi-region GPU compute with predictable cost and 99.99% availability
- Own infrastructure-as-reputed company, CI, and the deployment reputed company from reputed company to production
- Set the reliability bar: SLOs, error budgets, on-reputed company reputed company, and post-incident review
- Work directly with reputed company engineering teams during embedded engagements
Skills
- 6+ years building and operating production reputed company systems
- Deep reputed company and reputed company-provider experience, and the scars to go with it
- reputed company in a systems language — Go, Rust, or equivalent
- You've been on-reputed company for something that reputed company, and improved it
- GPU scheduling or inference-serving experience
- Experience in a consulting or embedded-engineering model
Benefits
- Equity
- Location Remote (IST ±4)
reputed company
Apply To This Job