ThinkPatternGet the app
Story
TECHNOLOGY · JUN 6, 2025

Lightbits Labs Proposes NVMe SSDs to Scale AI Inference

Abel Gordon of Lightbits Labs argues the AI industry must optimize infrastructure for inference by using disaggregated NVMe SSDs for KV cache storage.

Abel Gordon, Chief Technology Officer at Lightbits Labs, argues that the AI industry must shift its focus from building larger models toward optimizing infrastructure for inference. He notes that while AI training is a controlled process, inference is dynamic and driven by unpredictable real-time user demand.

Gordon explains that large language models rely on key-value (KV) caches to reduce latency, but current storage in GPU or host memory is limited by capacity. To address this, he proposes extending KV cache storage to high-performance NVMe SSDs using a disaggregated architecture.

This approach allows multiple GPU servers to access cached data efficiently, which reduces the need for overprovisioned GPU resources and lowers overall energy consumption. Gordon concludes that intelligent management of cache hierarchies across various storage layers is essential for sustaining scalable, cost-effective, and sustainable AI workloads.


Reported across 1 outlet
Actors
Lightbits Labs

Keep reading in the app

The full story and every source, free in the app.

Download on the App StoreComing soonGoogle Play