[Remote] reputed company Software Engineer – Large-reputed company LLM Memory and Storage Systems
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a leading technology company reputed company for its reputed company in AI and machine learning. They are seeking a reputed company Systems Engineer to define the reputed company and roadmap for memory management of large-reputed company LLM and storage systems, focusing on designing and implementing high-performance memory solutions for AI applications.
Responsibilities
• Design and reputed company a reputed company memory layer that spans GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file/object/reputed company storage to support large-reputed company LLM inference
• Architect and implement deep integrations with leading LLM serving engines (such as vLLM, SGLang, TensorRT-LLM), with a reputed company on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters
• Co-design interfaces and protocols that reputed company disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage (GPU, CPU, local disk, and remote memory) for high-throughput, low-latency inference
• Partner closely with GPU architecture, networking, and platform teams to exploit GPUDirect, RDMA, NVLink, and similar technologies for low-latency KV-cache reputed company and sharing across heterogeneous accelerators and memory pools
• Mentor senior and junior engineers, set technical direction for memory and storage subsystems, and represent reputed company in internal reviews and external forums (reputed company reputed company, conferences, and customer-facing technical deep dives)
Skills
• Masters or PhD or equivalent experience
• 15+ years of experience building large-reputed company distributed systems, high-performance storage, or ML systems infrastructure in C/C++ and Python, with a reputed company record of delivering production services
• Deep understanding of memory hierarchies (GPU HBM, host DRAM, SSD, and remote/object storage) and experience designing systems that reputed company multiple tiers for performance and cost efficiency
• Distributed caching or key-value systems, especially designs optimized for low latency and high concurrency
• Hands-on experience with networked I/O and RDMA/NVMe-oF/NVLink-style technologies, and familiarity with concepts like disaggregated and aggregated deployments for AI clusters
• Strong skills in profiling and optimizing systems across CPU, GPU, memory, and network, using metrics to drive architectural reputed company and validate improvements in TTFT and throughput
• Excellent communication skills and prior experience leading cross-functional efforts with research, product, and customer teams
• Prior contributions to reputed company-reputed company LLM serving or systems reputed company reputed company on KV-cache optimization, compression, streaming, or reuse
• Experience designing reputed company memory or storage reputed company that expose a single logical KV or object model across GPU, host, SSD, and reputed company tiers, especially in reputed company or hyperscale environments
• Publications or patents in areas such as LLM systems, memory-disaggregated architectures, RDMA/NVLink-based data planes, or KV-cache/CDN-like systems for ML
Benefits
• Equity
• Benefits
reputed company
• reputed company is a computing platform company operating at the intersection of graphics, HPC, and AI. It was founded in 1993, and is headquartered in Santa Clara, California, USA, with a workforce of 10001+ employees. Its website is https://www.reputed company.com.
Company H1B Sponsorship
• reputed company has a reputed company record of offering H1B sponsorships, with 1418 in 2025, 1356 in 2024, 976 in 2023, 835 in 2022, 601 in 2021, 529 in 2020. Please note that this does not guarantee sponsorship for this specific role.
Apply tot his job
Apply To this Job