Author | Wei Rong Editor | Bao Yonggang
The KV Cache for inference has grown so large that it now requires dedicated management.
This has become an increasingly clear consensus in the industry. Recently, Huawei launched the OceanStor M900 at the Huawei Connect conference, and NVIDIA has also introduced the CMX Context Memory Storage.
These AI memory storage products will hierarchically manage and store the KV Cache generated during the inference process that may be reused, reducing corresponding redundant computations when subsequent requests hit the cache, which will directly save computing power overhead. An industry insider explicitly stated that even as the cache hit rate continues to rise, improving it further from 90% to 95% remains highly valuable.
However, at a time when storage prices are "more expensive than gold," what exactly is the cost of independently storing some potentially reusable inference caches?
1、From DRAM to SSD, the Conversion Is Not 1:1
"The core logic of externalizing the KV Cache is 'trading storage for computing'," Zhang Heng, a scholar who has long studied storage, told Leiphone.
Trading storage for computing is essentially finding a balance between computing costs and storage costs. For historical KVs that have already been computed, should they be directly discarded and recomputed next time, or saved for direct reuse when subsequent requests hit the cache? The answer given by the industry is clearly the latter.
Zhang Heng gave an example: assuming the KV Cache hit rate increases from 90% to 95%, the portion that needs to be recomputed would drop from 10% to 5%. "This is equivalent to halving the computing power overhead for the recomputed portion," Zhang said.
However, a continued increase in the hit rate also means more historical KVs need to be saved. Especially in the high hit rate range, to cover the last few percentage points of long-tail data, the required storage capacity may increase significantly.
The first account to calculate is whether these KV Caches are worth storing.
According to Li Hui, head of a leading storage vendor, the data lifecycle varies greatly across different applications. In some consumer-end applications, the KV Cache might be cleared after just a few minutes or an hour or two.
Meanwhile, some enterprise-end applications place more emphasis on context continuity, and the data lifecycle of the KV Cache may also be longer. For such businesses, keeping historical KVs across requests for a longer time for subsequent reuse holds higher value.
The second account is where to store them. Li Hui believes that if DRAM were cheap enough and sufficiently supplied, putting everything in DRAM would naturally be the most convenient and offer the best performance—but the reality is that DRAM is both expensive and scarce, leaving room for SSDs.
However, this replacement is not simply swapping 1TB of DRAM for 1TB of SSD. On the one hand, Li Hui noted that from an overall cost perspective, the replacement between DRAM and SSD is not 1:1, but "1:N," and in some solutions, it could even reach 1:5 or 1:10. On the other hand, calculated at the same capacity, Li Hui estimates that the cost of DRAM and enterprise-grade SSDs can differ by dozens of times.
It is precisely this price difference that has gradually moved memory offloading from a "fallback measure" during past resource shortages into the initial design of new systems.
Although enterprises can use cheaper SSDs to replace part of the DRAM capacity, the scale of SSDs that need to be added will also be larger, and ultimately, capacity, performance, and procurement costs must still be calculated simultaneously. Zhang Heng told Leiphone that at this stage, most inference scenarios have not yet reached the point where they must add independent KV storage devices to operate.
Enterprises will usually first utilize existing DRAM, memory pools, and local SSDs before gradually considering external storage. "They will definitely max out the existing resources within the server first," Zhang said.
The emergence of Huawei's M900 and NVIDIA's CMX indicates that this KV tiering, which originally occurred mostly within the server, has begun to further move towards an independent shared storage layer.
2、KV Demand Emerges, but SSD Standards Fail to Keep Up
Independent KV storage devices have emerged, but the SSDs that truly host these KV Caches have not yet formed a mature and unified product category.
Li Hui told Leiphone that the industry has now begun to discuss SSDs optimized for AI inference, especially for KV Cache workloads, but when it actually comes to procurement, most customers are still buying standard enterprise-grade SSDs.
"Everyone in the industry is talking about this trend now, but there is no standardized, large-scale product yet," Li said. Currently, the market is still dominated by traditional enterprise-grade SSDs from vendors such as Samsung, SanDisk, and Micron.
The reason is not that SSDs cannot be built, but that the industry has not yet fully figured out exactly what kind of workload the KV Cache will bring to SSDs.
Li Hui gave an example of a related solution he encountered, where the initial requirement for SSD write endurance was about 3 DWPD, which was later increased to 9 DWPD. DWPD stands for Drive Writes Per Day, indicating how many full-drive writes an SSD can endure per day. "It might change to 12 tomorrow, and then back to 5 the day after," Li said.
The reason why the indicators keep changing depends first on how long the KV Cache actually needs to be saved. If a batch of caches is only kept for a few minutes, the data needs to be updated frequently, the write pressure on the SSD per day will be higher, and the requirement for endurance will also rise accordingly.
However, if the data needs to be saved for several weeks, after the write frequency drops, write endurance is no longer the primary contradiction, and capacity requirements will become more prominent.
Li Hui mentioned that the data lifecycle in different businesses now can range from minutes and hours to days or even weeks. Under the condition that application forms and data lifecycles have not yet stabilized, it is difficult to determine in advance exactly how much capacity, how high endurance, and what performance indicators the SSD will need.
For storage vendors, the SSD itself is a product highly dependent on scale. "Without standardization, it is difficult for this market to achieve scale," Li admitted.
If different customers define their own set of SSD specifications based on their businesses, the products will easily remain at the customization stage and fail to form a sufficiently large unified market. Li Hui gave an example: if the top internet giants ultimately need completely different products, it will be difficult for storage vendors to achieve scale relying on a single specification.
Therefore, the industry still needs time to integrate different solutions, and a more unified product standard may only be formed after the actual workloads gradually converge. According to Li Hui's estimate, it may take another 6 to 12 months for the industry to form a relatively unified product consensus; even after the standards are determined, subsequent product development, verification, and optimization may take another year or so.
Li Hui also mentioned that in the North American market, some customers have already increased their SSD procurement due to AI inference and KV Cache; the Chinese market has also started to see additional orders, but at a relatively slower pace. Currently, the bulk of domestic SSD demand is still in training and conventional inference, while KV Cache-related demand is still in the solution stage.
"In the future, KV Cache may indeed become one of the important growth drivers for enterprise-grade SSDs, but this market is currently still in a stage of diverse and fragmented development," Li said.
3、Storage Pressure Relieved, but Data Movement Costs Remain
Since so much historical KV Cache has been moved away, have the workloads of HBM and DRAM been alleviated?
The most directly alleviated pressure at present is the load on DRAM. Zhang Heng told Leiphone that one of the goals of externalizing the KV Cache is to reduce the server's reliance on large-capacity DRAM.
"The core goal is actually to replace main memory, not video memory. Main memory is too expensive now, so it might as well buy less memory and more SSDs," Zhang joked. In such solutions, some historical KVs that originally needed to stay in DRAM for a long time can be further offloaded to SSDs, thereby reducing memory configuration.
In contrast, the externalization of KV Cache does not alleviate HBM pressure as directly. Zhang Heng believes that external storage mostly affects the time-to-first-token latency. The speed at which historical KVs are fetched back from external storage will affect how soon the model starts generating the first token; and once token generation begins, the relevant KVs are already back in the video memory, and the subsequent generation speed still relies on high-speed video memory.
Li Hui also mentioned that as the price difference and supply disparities between different storage tiers expand, tiering and offloading are entering the design of inference systems earlier.
However, putting historical KVs into larger-capacity storage only solves the capacity and cost issues. During the current inference process, model weights and data still need to be constantly moved.
SSDs solve the problem of having enough storage capacity and cost-effectiveness, but they cannot solve the problem of whether the data can be moved fast enough.
Gu Huai, founder of an AI chip company, told Leiphone that Compute-in-Memory (CIM) attempts to reduce exactly the data movement overhead of model weights and current data. In traditional inference architectures, model weights and intermediate data need to be frequently moved between storage units and computing units; CIM aims to perform as much computation as possible near the data, reducing the overhead caused by repeated data movement.
For current requests to generate tokens quickly, relevant data will be placed in DRAM, or SRAM closer to the computing units; if data such as model weights and KVs can be reused near the computing location, repeated data movement can be reduced.
Gu Huai believes that this advantage is particularly evident in the Prefill stage; in high-concurrency scenarios, the same set of model parameters can serve multiple requests. Gu Huai took 100-way concurrency as an example: the same set of parameters can be reused by multiple requests, so there is no need to repeatedly move them for each request. "The higher the concurrency, the better," Gu said.
In the Decode stage, the KV caches used by different requests are not exactly the same, and less data can be reused across requests. Therefore, the benefits of CIM in this stage are not as obvious as in the Prefill stage.
In Gu Huai's view, if CIM enters servers in the future, a direct effect will be to reduce the time-to-first-token latency and support higher concurrency.
However, CIM cannot replace the storage requirements for historical KV Caches. A large amount of KV generated by historical conversations still needs to be saved in large-capacity media such as SSDs, and when subsequent requests hit the cache again, the required data is fetched back into the current computing path. "The KV Cache of historical conversations still needs to be stored in SSDs; otherwise, the storage volume would be too massive," Gu said.
So, returning to that question again, what exactly is the cost of storing those potentially reusable inference caches?
It is not simply a matter of spending a bit more on storage costs to save an equivalent amount of computing costs. How long the KV Cache is kept, how many times it can be reused, what medium is used, and how much data movement cost is incurred during retrieval all affect the final choice.
Large models remember the past, and infrastructure is also beginning to pay for "memory."