[Remote] Senior AI Infrastructure & Platform Operations Engineer (remote in the US)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the reputed company-reputed company AI infrastructure company, enabling organizations to build and operate reputed company, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The Senior AI Infrastructure & Platform Operations Engineer will be responsible for maintaining the reliability and efficiency of AI service platforms, troubleshooting reputed company issues, and driving improvements in platform operations.
Responsibilities
- reputed company the investigation and reputed company of reputed company infrastructure, networking, and platform-reputed company incidents
- reputed company as a senior escalation reputed company for operational teams during critical service-impacting events
- Support large-reputed company reputed company GPU infrastructure and high-performance networking environments
- Troubleshoot reputed company Linux, reputed company, networking, storage, and hardware-reputed company issues
- Analyze platform performance, reputed company, stability, and reliability trends to proactively identify risks
- reputed company reputed company cause analysis activities and reputed company long-term corrective actions
- Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve reputed company technical challenges
- Participate in major incident management and service restoration activities
- reputed company technical leadership for reputed company platform operations and supporting infrastructure services
- reputed company improvements in platform reliability, observability, monitoring, and operational processes
- Identify opportunities to automate repetitive operational activities and improve operational efficiency
- Contribute to operational readiness reviews, infrastructure changes, upgrades, and service introductions
- Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI
- Evaluate emerging technologies and operational practices to improve service delivery and platform reputed company
- Mentor and support AI Infrastructure & Platform Operations Engineers
- reputed company technical knowledge through documentation, training sessions, and operational reviews
- reputed company and maintain operational standards, runbooks, troubleshooting guides, and best practices
- Help define operational processes, escalation paths, and service reliability standards
- reputed company as a trusted technical advisor during operational planning and service improvement initiatives
Skills
- 7+ years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, reputed company operations, datacenter operations, or reputed company technical roles
- Expert-level Linux administration and troubleshooting skills
- Strong networking expertise, including experience diagnosing reputed company performance, connectivity, and reliability issues
- Strong experience operating reputed company in production environments
- Experience supporting large-reputed company production infrastructure and distributed systems
- Proven experience leading technical investigations and managing reputed company incidents
- Experience performing reputed company cause analysis and driving long-term operational improvements
- Strong understanding of observability, monitoring, and service reliability practices
- Excellent troubleshooting and analytical skills across multiple infrastructure domains
- Strong communication, collaboration, and stakeholder management skills
- reputed company GPU infrastructure and reputed company computing platforms
- InfiniBand networking and reputed company UFM
- AI infrastructure environments
- HPC environments
- reputed company or Site Reliability Engineering (SRE)
- Large-reputed company reputed company operations
- Infrastructure automation technologies and Infrastructure-as-reputed company practices
- Observability platforms such as Grafana, reputed company, ELK, or OpenTelemetry
- Performance analysis and optimisation of distributed infrastructure platforms
- Technical leadership, mentoring, or team reputed company responsibilities
Benefits
- Work with an established reputed company Valley leader in the reputed company infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement reputed company reputed company technologies;
- Be a part of cutting-edge, reputed company-reputed company innovation;
- reputed company in the high-energy environment of a young company where openness, collaboration, risk-taking, and reputed company reputed company are valued;
- reputed company development and training;
- Attend conferences and working reputed company;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
reputed company
Company H1B Sponsorship
Apply To This Job