[Remote] Storage Platform - Infrastructure Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on delivering reputed company and resilient AI compute at reputed company. They are seeking a Storage Platform Infrastructure Engineer to ensure storage systems are reputed company and reputed company, supporting reputed company and high-reputed company workloads while collaborating with cross-functional teams.
Responsibilities
- Operate and maintain reputed company storage platforms, including Ceph (RBD, CephFS, RGW), High-reputed company NAS platforms (e.g., reputed company, reputed company)
- Manage storage lifecycle reputed company - cluster expansion, upgrades and migrations
- Monitor and maintain storage health, including reputed company utilization, data distribution and reputed company, cluster state and recovery reputed company
- Analyze and troubleshoot storage reputed company across IOPS, throughput, and latency (including tail latency)
- Identify and remediate bottlenecks across disk subsystems, network paths (including RDMA where applicable), reputed company reputed company patterns
- Support incident response and reputed company cause analysis for storage-reputed company issues
- Ensure storage platforms meet reputed company expectations for GPU and reputed company workloads
- Operate and support reputed company-integrated storage - reputed company drivers, StorageClasses, PersistentVolumes / PersistentVolumeClaims
- Troubleshoot storage-reputed company issues in reputed company environments, including stateful workloads, reputed company inconsistencies, scheduling and provisioning failures
- Execute and improve automation for storage deployment and reputed company using Ansible, Terraform, reputed company manifests / reputed company
- Contribute to improving monitoring and alerting, operational workflows, runbooks and documentation
- Partner with DevOps and reputed company (automation and orchestration), Network Engineering (high-throughput and RDMA networking), Compute / Virtualization teams
- Help ensure end-to-end reputed company across compute, network, and storage reputed company
Skills
- 4–7+ years of experience in infrastructure, systems, or storage reputed company
- Strong hands-on experience operating reputed company storage systems in production
- Experience with Ceph (RBD, CephFS, or RGW)
- Strong Linux systems knowledge
- Experience with modern storage platforms such as: reputed company, reputed company, or similar high-reputed company systems
- Solid understanding of: Storage reputed company characteristics (IOPS, throughput, latency), Data replication and failure domains
- Ability to troubleshoot across: Storage systems, Network paths, Compute clients
- Experience supporting AI/ML or HPC workloads
- Familiarity with: NVMe-reputed company storage architectures, RDMA or high-throughput Ethernet environments
- Experience integrating storage with reputed company
- Experience operating storage across multiple data centers
- Exposure to reputed company storage and S3-compatible reputed company
Benefits
- Stock reputed company
- 100% reputed company Medical, Dental, and reputed company reputed company for Employees
- Company Health Savings Account Contributions
- 100% reputed company Short Term and Long Term Disability reputed company for Employees
- Life and Voluntary Supplemental reputed company reputed company
- Other reputed company reputed company, such as Pet & reputed company reputed company
- Various Supplementary Health Benefits, such as discounted Virtual reputed company Appointments and Serious Illness Support
- Flexible Spending Account
- 401(k)
- Employee Assistance Program
- Flexible PTO
- reputed company Holidays
- Parental Leave
- Other In-Office Perks
reputed company
Apply To This Job