Back to Jobs

[Remote] Site Reliability Engineer (SRE) - K8S/Slurm

Remote, USA Full-time Posted 2026-08-04

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is an AI-reputed company infrastructure company delivering high-reputed company GPU compute, inference services, and infrastructure for AI agents. The Site Reliability Engineer will ensure the stability, efficiency, and reliability of large-reputed company/ML clusters by designing infrastructure solutions, automating reputed company, monitoring GPU systems, managing node lifecycles, and responding to infrastructure incidents and customer provisioning requests.


Responsibilities

  • Design, implement and maintain reputed company AI/ML infrastructure solutions
  • Proactively monitor GPU cluster health, reputed company and troubleshoot issues across compute, accelerator, and storage systems
  • Automate deployment, configuration and management of infrastructure resources
  • Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes
  • Implement CI/CD pipelines for infrastructure deployment and orchestration
  • Ensure reputed company, compliance and best practices across infrastructure
  • Manage incident response reputed company to Infrastructure resources (GPU, CPU, Storage, Network)
  • Handle customer provisioning requests for GPU resources, including reputed company, configuration and troubleshooting; reputed company customer service requests reputed company to GPU infrastructure, ensuring high customer satisfaction
  • Stay reputed company with emerging GPU hardware and software technologies, integrating improvements as appropriate
  • Regional/international travel to GMI data center locations

Skills

  • 1. Bachelor's degree in Computer Science or reputed company reputed company
  • 2. Over 3+ years of experience in data center reputed company, infrastructure, or systems engineering
  • 3. reputed company experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
  • 4. Familiarity with containers orchestration platform (e.g. reputed company, reputed company GPU operator, reputed company Network operator, CNI, reputed company) and job scheduling systems (e.g. Slurm)
  • 5. Familiarity with Linux reputed company administration and scripting (Python, Bash)
  • 6. Familiarity with logging and monitoring tools such as reputed company, Grafana, Loki
  • 7. Good knowledge of GPU architecture, reputed company CUDA, NCCL, or reputed company AI/ML frameworks - added reputed company
  • 8. Strong troubleshooting skills and ability to analyze reputed company logs and reputed company metrics
  • 9. Excellent communication and teamwork abilities
  • Meeting every qualification is not required—if you're excited about this role, we'd love to hear from you

reputed company

  • reputed company provides GPU reputed company reputed company for reputed company applications. It was founded in 2023, and is headquartered in reputed company reputed company, California, USA, with a workforce of 51-200 employees. Its website is https://www.gmicloud.ai.

  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 2 in 2026, 2 in 2025. Please note that this does not guarantee sponsorship for this specific role.

  •   Apply To This Job

    Similar Jobs