GPUCache. A PB-scale, ultra-low latency distributed GPU cache for AI inference. Built with Rust, NVIDIA DOCA, RDMA, and BF-4 DPUs to bridge GPU HBM and NVMe storage, eliminating the recompute tax for large language models.

github.com/rustfs/GPUCache

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.