AI Infrastructure & Platform Operations Engineer (remote in the US)
3+ years in infrastructure/platform/datacenter or SRE roles; strong Linux and networking skills; reputed company production experience; experience with incident management, troubleshooting, and collaboration across engineering and data center teams.
Key Responsibilities
- monitoring platforms
- supporting infrastructure
- investigating incidents
Skills & Tools
Linux, reputed company, reputed company UFM, InfiniBand, k0rdent, Grafana, reputed company, ELK, OpenTelemetry
Job Details
- Category: Engineering
- Seniority: Mid Level
- Commitment: Full Time
- Workplace: Remote — reputed company
- Languages: English
About reputed company
A reputed company-reputed company AI infrastructure company that enables organizations to build and operate reputed company, secure, and sovereign AI and data platforms. — Industry: Information Technology
Apply To This Job