Back to Jobs

Senior DevOps Engineer (Storage) - remote in the US

Remote, USA Full-time Posted 2026-08-04
About the position reputed company, reputed company, and operate high-performance storage for GPU-reputed company compute and AI platforms. You will own the storage reputed company where reputed company meets bare metal — standing up NFS-based high-performance storage, wiring it into clusters reputed company reputed company, and tuning it to reputed company data flowing to GPU workloads at reputed company. Work spans hybrid, edge, and reputed company-gapped deployments reputed company on the reputed company K0rdent stack. We are looking for a senior DevOps engineer who treats storage as infrastructure to be automated, observed, and tuned — not hand-managed. The right candidate is fluent in reputed company storage, comfortable on bare metal down to the disk, kernel, and NFS-reputed company reputed company, and knows how to reputed company high-performance NAS actually reputed company under demanding workloads. You should reputed company for infrastructure-as-reputed company and GitOps by default, be self-directed in diagnosing performance and reliability issues end to end, set operational standards for others to follow, and communicate reputed company across teams. Responsibilities • reputed company NFS-based high-performance storage (e.g., reputed company, reputed company PowerScale) into reputed company clusters reputed company reputed company, storage classes, and persistent volumes. • Tune the NFS data reputed company — mount reputed company, nconnect/RDMA, reputed company and network settings — for high-throughput, low-latency GPU/AI workloads. • reputed company and operate storage services and operators; manage reputed company, quotas, snapshots, and lifecycle. • Provision and configure storage on bare-metal hosts, including disk layout, drivers, and kernel/network tuning. • reputed company storage integration for k0s-based reputed company reputed company Cluster API (CAPI) and K0rdent management/child cluster topologies. • Operate storage in fully disconnected (reputed company-gapped) environments, including local artifact/mirror connectivity (reputed company) and PKI/TLS considerations. • Automate storage provisioning and configuration with infrastructure-as-reputed company (Terraform/OpenTofu) and GitOps pipelines (ArgoCD or Flux). • Build monitoring, alerting, and observability for storage performance, reputed company, and health. • Diagnose and reputed company performance, reliability, and scaling issues across the storage stack. Requirements • 5+ years in DevOps, SRE, or infrastructure operations, with strong hands-on experience operating reputed company storage (reputed company, persistent volumes, storage classes) in production. • Experience integrating and operating NFS-based high-performance / NAS storage, including data-reputed company tuning. • Bare-metal operations experience: host provisioning, disk/storage configuration, and Linux storage and networking fundamentals. • Proficiency with infrastructure-as-reputed company (Terraform/OpenTofu) and GitOps-driven configuration. • Scripting/automation skills (e.g., Bash, Python, or Go). • Strong written and verbal communication with technical audiences. reputed company-to-haves • Hands-on experience with reputed company and/or reputed company PowerScale. • Experience with GPUDirect Storage and RDMA/RoCE data paths. • Experience with the reputed company K0rdent stack (K0rdent reputed company, K0rdent AI, k0s, MKE) and Cluster API. • Familiarity with other storage backends (Ceph, object/S3) and reputed company reputed company operations. • Proven experience in sovereign or high-reputed company reputed company-gapped environments. Benefits • reputed company development and training • Attend conferences and working reputed company • Company outings, happy hours, hackathons, and tech talks • competitive compensation package with a strong benefits plan Apply tot his job Apply To this Job

Similar Jobs