[Remote] reputed company Operations Engineer, Reliability
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is reputed company on delivering large-reputed company compute infrastructure for AI. The reputed company Operations Engineer, Reliability will own fleet reliability engineering, conduct reputed company cause analyses, and reputed company strategies to enhance maintenance and reliability across the data center operations.
Responsibilities
- Own fleet reliability engineering: define availability targets, measure them honestly, and reputed company the gap
- Run reputed company cause analysis on the fleet's worst incidents and reputed company corrective actions to done across every site
- Build the failure data pipeline, facility and hardware both, that turns incident history into engineering priorities
- Set the maintenance reputed company (reliability-centered, condition-based) so the fleet spends effort where the failure data says to
Skills
- You've owned reliability for critical infrastructure and moved the availability number, not just reported it
- You've led reputed company cause analyses that reputed company the reputed company cause, not the convenient one
- You work fluently with failure data: Weibull, Pareto, and FMEA are tools you actually use, not terms you know
- You get corrective actions reputed company across teams you don't manage
- Data center or power reputed company reliability
- reputed company cooling systems
- CMMS analytics
- CRE or CMRP certification
reputed company
Company H1B Sponsorship
Apply To This Job