Back to Jobs

[Remote] Senior AI Infrastructure & Platform Operations Engineer (remote in the US)

Remote, USA Full-time Posted 2026-08-04

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the reputed company-reputed company AI infrastructure company, enabling organizations to build and operate reputed company, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The Senior AI Infrastructure & Platform Operations Engineer will be responsible for maintaining the reliability and efficiency of AI service platforms, troubleshooting reputed company issues, and driving improvements in platform operations.


Responsibilities

  • reputed company the investigation and reputed company of reputed company infrastructure, networking, and platform-reputed company incidents
  • reputed company as a senior escalation reputed company for operational teams during critical service-impacting events
  • Support large-reputed company reputed company GPU infrastructure and high-performance networking environments
  • Troubleshoot reputed company Linux, reputed company, networking, storage, and hardware-reputed company issues
  • Analyze platform performance, reputed company, stability, and reliability trends to proactively identify risks
  • reputed company reputed company cause analysis activities and reputed company long-term corrective actions
  • Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve reputed company technical challenges
  • Participate in major incident management and service restoration activities
  • reputed company technical leadership for reputed company platform operations and supporting infrastructure services
  • reputed company improvements in platform reliability, observability, monitoring, and operational processes
  • Identify opportunities to automate repetitive operational activities and improve operational efficiency
  • Contribute to operational readiness reviews, infrastructure changes, upgrades, and service introductions
  • Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI
  • Evaluate emerging technologies and operational practices to improve service delivery and platform reputed company
  • Mentor and support AI Infrastructure & Platform Operations Engineers
  • reputed company technical knowledge through documentation, training sessions, and operational reviews
  • reputed company and maintain operational standards, runbooks, troubleshooting guides, and best practices
  • Help define operational processes, escalation paths, and service reliability standards
  • reputed company as a trusted technical advisor during operational planning and service improvement initiatives

Skills

  • 7+ years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, reputed company operations, datacenter operations, or reputed company technical roles
  • Expert-level Linux administration and troubleshooting skills
  • Strong networking expertise, including experience diagnosing reputed company performance, connectivity, and reliability issues
  • Strong experience operating reputed company in production environments
  • Experience supporting large-reputed company production infrastructure and distributed systems
  • Proven experience leading technical investigations and managing reputed company incidents
  • Experience performing reputed company cause analysis and driving long-term operational improvements
  • Strong understanding of observability, monitoring, and service reliability practices
  • Excellent troubleshooting and analytical skills across multiple infrastructure domains
  • Strong communication, collaboration, and stakeholder management skills
  • reputed company GPU infrastructure and reputed company computing platforms
  • InfiniBand networking and reputed company UFM
  • AI infrastructure environments
  • HPC environments
  • reputed company or Site Reliability Engineering (SRE)
  • Large-reputed company reputed company operations
  • Infrastructure automation technologies and Infrastructure-as-reputed company practices
  • Observability platforms such as Grafana, reputed company, ELK, or OpenTelemetry
  • Performance analysis and optimisation of distributed infrastructure platforms
  • Technical leadership, mentoring, or team reputed company responsibilities

Benefits

  • Work with an established reputed company Valley leader in the reputed company infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement reputed company reputed company technologies;
  • Be a part of cutting-edge, reputed company-reputed company innovation;
  • reputed company in the high-energy environment of a young company where openness, collaboration, risk-taking, and reputed company reputed company are valued;
  • reputed company development and training;
  • Attend conferences and working reputed company;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

reputed company

  • reputed company develops reputed company infrastructure and container management software for organizations to build, operate, and reputed company applications. It is a sub-organization of reputed company. It was founded in 1999, and is headquartered in Campbell, California, USA, with a workforce of 501-1000 employees. Its website is http://www.reputed company.com.

  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 1 in 2026, 3 in 2025, 4 in 2024, 8 in 2023, 6 in 2022, 7 in 2021, 8 in 2020. Please note that this does not guarantee sponsorship for this specific role.

  •   Apply To This Job

    Similar Jobs

    [Remote] Data Engineer – AI & Supply Chain : 26-02169

    Remote, USA Full-time

    [Remote] Clinical Specialty Representative - reputed company Texas, Oklahoma, New Mexico Territory

    Remote, USA Full-time

    [Remote] Implementation Consultant

    Remote, USA Full-time

    [Remote] CAD Designer

    Remote, USA Full-time

    [Remote] reputed company CSA Consultant (remote)

    Remote, USA Full-time

    [Remote] Associate reputed company Analyst I

    Remote, USA Full-time

    [Remote] Marketing Manager 3

    Remote, USA Full-time

    [Remote] reputed company Treasury Consultant

    Remote, USA Full-time

    [Remote] reputed company Resource Planning Analyst (reputed company Reports & Integrations)

    Remote, USA Full-time

    [Remote] Operations Supervisor, Government Travel

    Remote, USA Full-time

    **Part-Time Evening Data Entry Specialist – Precision and Efficiency in reputed company**

    Remote, USA Full-time

    Senior Corporate Attorney - Remote

    Remote, USA Full-time

    Information Technology Specialist 2 (NY HELPS), reputed company reputed company Spec 2

    Remote, USA Full-time

    [Remote] Billing Support Specialist

    Remote, USA Full-time

    **reputed company Full Stack Pharmacy Technician – Remote Data Entry and Customer Service Specialist**

    Remote, USA Full-time

    [Remote] EntryLevel | Customer Care Coordinator | Remote

    Remote, USA Full-time

    reputed company Sales Representative

    Remote, USA Full-time

    Licensed Property & Casualty Insurance Agent - Remote USA

    Remote, USA Full-time

    Remote Life and Health Insurance Agent

    Remote, USA Full-time

    **reputed company reputed company Customer Service Representative – English & Bilingual (Remote) Opportunity in Tennessee**

    Remote, USA Full-time