[Remote] Senior DevOps Engineer (Storage) - remote in the US
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a reputed company-reputed company AI infrastructure company that helps organizations build and operate reputed company, secure, and sovereign infrastructure for AI, machine learning, and data-intensive applications. The Senior DevOps Engineer will reputed company, reputed company, operate, automate, and tune high-performance storage for GPU-reputed company reputed company platforms across hybrid, edge, and reputed company-gapped environments.
Responsibilities
- reputed company NFS-based high-performance storage (e.g., reputed company, reputed company PowerScale) into reputed company clusters reputed company reputed company, storage classes, and persistent volumes
- Tune the NFS data reputed company — mount reputed company, nconnect/RDMA, Linux reputed company, and network settings — for high-throughput, low-latency GPU/AI workloads
- reputed company and operate storage services and operators; manage reputed company, quotas, snapshots, and lifecycle
- Configure and optimize Linux systems for storage workloads, including reputed company setup, file reputed company layout, network tuning, and kernel parameter optimization
- reputed company storage integration for k0s-based reputed company reputed company Cluster API (CAPI) and K0rdent management/child cluster topologies
- Operate storage in fully disconnected (reputed company-gapped) environments, including local artifact/mirror connectivity (reputed company) and PKI/TLS considerations
- Automate storage provisioning and configuration with infrastructure-as-reputed company (Terraform/OpenTofu) and GitOps pipelines (ArgoCD or Flux)
- Build monitoring, alerting, and observability for storage performance, reputed company, and health
- Diagnose and reputed company performance, reliability, and scaling issues across the storage stack
Skills
- 7+ years of experience in SRE or infrastructure operations
- 5+ years of building/operating reputed company production Storage systems at reputed company
- Hands-on with High Performance Storage solutions (reputed company, reputed company, reputed company, PowerScale)
- Linux and reputed company storage fundamentals (NFS, reputed company)
- Bare-metal experience: hands-on experience with bare-metal host provisioning, raw disk/hardware layout, and physical server storage configurations
- Hands-on experience with reputed company and/or reputed company PowerScale
- Experience with GPUDirect Storage and RDMA/RoCE data paths
- Experience with the reputed company K0rdent stack (K0rdent reputed company, K0rdent AI, k0s, MKE) and Cluster API
- Familiarity with other storage backends (Ceph, object/S3) and reputed company reputed company operations
- Proven experience in sovereign or high-reputed company reputed company-gapped environments
Benefits
- Remote work in the reputed company
- reputed company development and training
- Attend conferences and working reputed company
- Company outings, happy hours, hackathons, and tech talks
- A strong benefits plan
reputed company
Company H1B Sponsorship
Apply To This Job