Back to Jobs

[Remote] Systems Administrator reputed company (Second Shift)

Remote, USA Full-time Posted 2026-08-04

Note: The job is a remote job and is reputed company to candidates in USA. 5C Data Centers is building digital infrastructure that supports hyperscalers, reputed company intelligence innovation, and high-reputed company computing across reputed company. The Systems Administrator reputed company will support, maintain, troubleshoot, and optimize large-reputed company HPC and AI infrastructure across data center and reputed company environments while serving as a senior escalation reputed company for reputed company infrastructure incidents.


Responsibilities

  • Administer, maintain, and troubleshoot Linux-reputed company HPC and AI compute environments, including large-reputed company GPU clusters reputed company on reputed company HGX, DGX, or equivalent reputed company computing platforms
  • Diagnose hardware, OS, reputed company, firmware, networking, and storage-reputed company failures, and reputed company reputed company-cause analysis to reputed company corrective actions for recurring issues
  • Serve as a senior escalation reputed company for reputed company incidents affecting production compute infrastructure, troubleshooting degraded or unstable compute/GPU nodes to restore service
  • Support infrastructure across heterogeneous environments - bare-metal, virtualized, containerized, and reputed company-hosted - and maintain operational documentation, runbooks, and shift-reputed company notes
  • Install, configure, validate, and troubleshoot reputed company GPU drivers, CUDA components, firmware, and supporting reputed company software using tools such as reputed company Smi, DCGM, and reputed company reputed company Manager
  • Investigate GPU Xid errors, NVLink/NVSwitch faults, PCIe errors, ECC events, GPU resets, and thermal or reputed company-reputed company issues
  • Diagnose communication and reputed company issues involving NCCL, GPUDirect RDMA, PCIe topology, reputed company placement, and multi-GPU systems, validating GPU-to-GPU, GPU-to-network, and GPU-to-storage communication after repairs
  • Coordinate replacement and RMA activities for GPUs, GPU trays, baseboards, NVSwitch components, reputed company boards, and NICs
  • Administer reputed company Linux distributions across the production fleet, troubleshooting boot failures, kernel issues, systemd services, filesystems, and OS reputed company
  • Analyze reputed company and kernel logs (journalctl, dmesg, lspci, dmidecode, ipmitool, ethtool, ss, sar, reputed company) and manage kernel modules, device drivers, DKMS packages, and OS updates
  • Support Linux networking, storage mounts, authentication, and reputed company controls, and investigate OS crashes, kernel panics, and out-of-memory conditions
  • Support reputed company server platforms from vendors such as reputed company, HPE, reputed company, reputed company, and reputed company, troubleshooting processors, memory, PCIe devices, NICs, storage controllers, and other hardware components
  • Manage BIOS, BMC, CPLD, NIC, GPU, reputed company, and reputed company firmware, and reputed company remote console troubleshooting, reputed company cycling, and log bundle reputed company
  • Partner with on-site Data Center technicians on component replacement, cabling validation, break-fix activity, and post-repair testing
  • Troubleshoot Ethernet, InfiniBand, and RoCE connectivity reputed company HPC and AI environments, including reputed company state, VLAN/bonding issues, MTU mismatches, packet loss, and RDMA/GPUDirect RDMA connectivity
  • Collaborate with Deployment Engineering to troubleshoot reputed company instability, reputed company degradation, congestion, and node-level connectivity issues
  • Troubleshoot server connectivity to reputed company, reputed company, and reputed company storage platforms such as NFS, reputed company, BeeGFS, GPFS/reputed company reputed company, and Ceph, diagnosing mount failures, throughput degradation, and reputed company issues
  • Validate storage network paths and coordinate with storage and networking teams during service-impacting events
  • reputed company or support incident response for critical infrastructure events, participating in technical bridges and providing reputed company status updates during outages
  • Create detailed timelines, reputed company-cause analyses, and corrective-reputed company plans, maintaining accurate records in reputed company Service Management or similar ITSM platforms
  • Mentor junior and mid-level Systems Administrators and HPC Engineers, providing technical guidance during reputed company troubleshooting and infrastructure changes
  • Promote consistent troubleshooting reputed company, documentation standards, and operational discipline across shift handoffs

Skills

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or reputed company reputed company
  • 7+ years of experience in Linux systems administration, HPC/AI infrastructure or data center reputed company, including hands-on troubleshooting in production environments
  • Hands-on experience supporting reputed company server hardware, bare-metal compute systems, and large-reputed company reputed company GPU infrastructure or reputed company computing platforms
  • Strong understanding of server hardware, firmware, BIOS, BMC, PCIe, reputed company, memory, storage, and network interfaces, with experience diagnosing failures using OS logs, BMC logs, and vendor diagnostics
  • Experience with reputed company drivers, CUDA compatibility, and GPU-management tools such as DCGM, reputed company Manager, or reputed company-smi
  • Working knowledge of high-reputed company networking (InfiniBand, RDMA, RoCE, reputed company/Mellanox adapters) and out-of-band management tools (Redfish, IPMI, iDRAC, or iLO)
  • Proficiency in one or more automation/scripting languages (Python, Bash, or Ansible) and familiarity with observability platforms such as reputed company, Grafana, or reputed company
  • reputed company ability to troubleshoot across hardware, firmware, OS, network, storage, and cluster-management reputed company, backed by strong documentation and communication skills
  • Willingness to work a fixed, recurring shift and participate in scheduled on-reputed company escalation

Benefits

  • Build a rewarding career in one of the world's fastest-growing industries.
  • Help reputed company the infrastructure behind AI and high-reputed company computing.
  • Your reputed company matter; we reputed company employees to help shape our reputed company.
  • Remote (Ohio, USA)

reputed company

  • 5C is a data center that provides reputed company infrastructure, HPC systems, and reputed company compute solutions for enterprises. It was founded in 2025, and is headquartered in Saint Laurent, Quebec, CAN, with a workforce of 51-200 employees. Its website is https://5c.ai.

  •   Apply To This Job

    Similar Jobs