In recent years, the competition among smartphone manufacturers in AI has basically revolved around a few familiar aspects: model parameters, NPU computing power, and on-device deployment capabilities. Especially in the current era where on-device AI is becoming ubiquitous, AI has become one of the "must-have" features for smartphones.
However, as smartphone AI capabilities continue to strengthen, new problems have emerged: chips have greater computing power, and models need to process more data during operation. No matter how fast the processor is, it must first obtain the data. The larger the model, the more data needs to be stored and scheduled. Consequently, storage and memory naturally become crucial factors affecting the on-device AI experience.
At this year's Qualcomm Snapdragon Summit, the iSA (Intelligent Storage Agent) introduced by Longsys became one of the first products to be compatible with the sixth-generation Snapdragon 8 Elite, debuting alongside UFS 4.1. Previously, iSA had completed joint tuning on the AMD Ryzen AI Max+ 395 platform, and now it has been adapted for flagship smartphone platforms. It aims to solve the current dilemma in the smartphone market: as on-device models grow larger, the competition in smartphone AI can no longer focus solely on the CPU, GPU, and NPU. Where data is stored and how it is invoked will equally impact the actual user experience.
Fitting Large Models into Smartphones: Is Memory the First to Break Down?
In the past, AI on smartphones mostly executed single tasks, such as recognizing a photo or summarizing a piece of text. These tasks had short durations and required little data to be processed simultaneously. However, today's agents are different. They need to understand context, break down tasks, and invoke different models. Some tasks run continuously for a long time, and both model sizes and runtime are constantly increasing. The problem is that the DRAM capacity of smartphones cannot keep up with the growth rate of model sizes.
How severe has this problem become? In an exclusive interview with Lei Technology, Huang Qiang, Vice President of Longsys and General Manager of the Embedded Storage Business Unit, stated: "Currently, the DRAM in flagship smartphones is generally between 12GB and 16GB, which is fully sufficient for current daily needs. However, as 30B-class large models begin to enter smartphones, higher requirements are placed on capacity, speed, power consumption, cost, and energy efficiency. The internal space and thermal conditions of smartphones are limited, and unlike cloud servers, hardware resources cannot be continuously expanded."
Wang Shuchong, Vice President of Huiyiwei, a subsidiary of Longsys, further explained that the Android system itself occupies several gigabytes of memory. If some large apps are also running, the actual space left for AI is quite limited. Moreover, a 30B model without sufficient compression and optimization could occupy about 30GB just for its parameters.
In other words, even the most high-end flagship smartphones today cannot fit all model parameters into DRAM.
The 30B MoE (Mixture of Experts) model showcased at this Qualcomm Summit has undergone targeted optimizations for the model, inference, and storage. On the model side, it can be further compressed through pruning and quantization. The MoE architecture only activates a portion of expert parameters during each inference. Combined with on-device inference and storage optimizations, the resources required and scheduled during operation are ultimately controlled at the scale of several gigabytes. According to Longsys, this 30B model can already run smoothly on smartphone reference designs.
MoE, or the Mixture of Experts architecture, is now being adopted by an increasing number of large models. Its principle is easy to understand: a MoE model contains many "experts" responsible for different tasks. During each inference, not all experts are "dispatched"; instead, a subset is selected to participate in the computation based on the task's requirements.
Since only a portion of experts is used at a time, there is no need to keep the temporarily unactivated parameters in DRAM just to "occupy space." Longsys's iSA places the experts in the smartphone's larger-capacity NAND. When the model needs to invoke these "experts," the corresponding parameters are then moved into DRAM.
Alright, the space issue has a solution, but read and write speeds have become a new challenge.
Wang Shuchong noted that data from ordinary apps, photos, and files on smartphones is usually relatively continuous. NAND can often read sequentially, much like reading a book from front to back where the content is contiguous, resulting in fast reading speeds. However, MoE is different. Which expert needs to be invoked each time depends on the current task. This time it might invoke Expert A, and next time it could be Expert F. It is more like consulting a reference book, where you flip directly to the corresponding chapter based on what you need.
This reading method has also changed the requirements for storage. In the past, a lot of data could be read continuously in one go. Now, the model may frequently switch between different experts, requiring NAND to continuously read different data, which increases the importance of random reads. For on-device large models, it is not enough for UFS to just have high sequential read speeds; it must also be fast enough when facing these more fragmented data invocations.
More troublesome is the fact that no matter how fast NAND read and write speeds are, they still cannot match DRAM. If the experts needed by the model are already in DRAM, the smartphone's NPU can use them directly for computation. But if they are still in NAND, they must first be read and transferred before computation can begin, which will affect the output speed.
This is exactly the problem that iSA aims to solve.
Fitting Models into Smartphones Only Solves Half the Problem
The solution provided by Longsys is "intelligent prefetching": trying to transfer the corresponding data from NAND to DRAM before the NPU actually needs a specific expert.
But there is a prerequisite for reading ahead: iSA must know which experts the model might use next.
iSA references the model's own operational patterns while combining the current context. During the initial loading, it can first refer to existing expert activation patterns. Once the inference stage begins, the user's input prompt, previous context, and the current inference stage all become reference information. The upper-level inference engine passes this information to iSA, which then judges which experts might be used next and prepares the corresponding data in advance.
For example, if a user was just querying the model for information about Maui, Hawaii, the probability of continuing to ask about local hotels and attractions in the next round is obviously higher than suddenly switching to the weather in Beijing. For the model, there is a contextual connection between the two rounds of dialogue, and the corresponding expert invocations follow a certain pattern. iSA can use this information to expand the prefetching scope and transfer potentially needed experts from NAND to DRAM in advance.
Of course, no matter how smart prefetching is, it cannot guess correctly every time. If the experts transferred to DRAM in advance end up not being used, the actually needed data still has to be read from NAND again. Wang Shuchong revealed that iSA's prefetching strategy can also be adjusted to be more aggressive or more conservative based on actual conditions: prefetching more increases the hit rate but brings additional I/O; prefetching less reduces the pressure on storage, but the waiting time will be longer when predictions fail. Since the expert activation patterns vary across different models, iSA currently still needs to be adapted to specific models and inference engines.
iSA solves the problem of when to fetch and which data to fetch, but the actual data reading and writing still rely on storage hardware. Longsys's UFS 4.1, which debuted alongside iSA, takes on this part of the work.
As mentioned earlier, data reading when MoE invokes experts is relatively fragmented, placing higher demands on random reads. Longsys's UFS 4.1, paired with Huiyiwei's self-developed controller, can optimize random read efficiency and access latency for such AI workloads. This allows the expert parameters selected in advance by iSA to be transferred from NAND to DRAM faster, minimizing the time the NPU spends waiting for data.
Furthermore, Huiyiwei's self-developed UFS controller, in addition to performing actual data reading and writing, is also responsible for data placement, space management, and wear leveling, which also helps with storage lifespan and long-term stability.
Besides MoE expert weights, there is another element that continuously consumes memory during the operation of large models: the KV Cache.
When large models engage in multi-turn dialogues, to reduce redundant computations, they retain some previously calculated information. The longer the conversation and the longer the context, the more space the KV Cache will occupy. In the past, when smartphone AI performed a single image recognition or text summarization, the task ended quickly. In the future, agents may continuously plan a trip, analyze a long document, or even work continuously across multiple apps, significantly increasing the demand for KV Cache.
Smartphones only have a dozen or so gigabytes of DRAM, which must also be shared with the system and other apps. Naturally, the KV Cache cannot be continuously piled into it. iSA will manage this part of the data hierarchically, moving some cache to larger-capacity storage while clearing caches that no longer need to be retained. In this way, both expert weights and KV Cache can be scheduled between DRAM and NAND based on actual usage, reserving the limited DRAM for data that is more urgently needed at the moment.
Of course, after entrusting more AI runtime data to NAND, another problem arises: will the lifespan of the flash memory be affected?
After all, if agents truly become a resident feature on smartphones in the future, model parameters and related data will be frequently accessed every day. The read and write pressure on the memory will increase, and over time, concerns about lifespan will naturally arise. In response, Huang Qiang stated that a dozen-gigabyte model is very large for DRAM, but when placed in 256GB or 512GB NAND, the proportion is much smaller, and the impact on lifespan is quite limited. Moreover, Huiyiwei's self-developed controller can work with iSA for lifespan management, space optimization, and wear leveling. If certain areas are written to too many times, data locations can be readjusted to avoid long-term concentrated erasing and writing.
However, agents have not yet truly become high-frequency resident applications on smartphones. The read and write methods, usage frequency, and workloads of different models also vary. If AI runs for several hours a day in the future or even stays in the background long-term, exactly how much the write volume, power consumption, and temperature of NAND will increase can only be known after long-term operation on real terminals.
Also requiring verification through mass-produced products is how much DRAM iSA can ultimately save, how much it can reduce the first-token latency, and how much it can improve the token generation speed. The expert activation patterns of different models will affect the prefetching effect. Changing the model, inference engine, or even the terminal environment could result in different final data.
From Demo to Mass Production: How Far Away is iSA?
The sixth-generation Snapdragon 8 Elite is not the first computing platform adapted for iSA. Previously, Longsys had completed joint tuning on the AMD Ryzen AI Max+ 395 platform, covering product forms such as AI PCs and AI BOXes. However, the situation is somewhat different on smartphones. Smartphones have smaller memory capacities and are more sensitive to power consumption and heat dissipation, so iSA naturally faces more constraints.
Of course, the biggest issue for iSA at present is still adaptation. The expert activation patterns of different large models vary, and iSA needs to be adjusted separately according to the characteristics of each model. This demo uses Step's model and the Wulianghuo inference engine, requiring Longsys to jointly optimize with both parties. When it comes to actual smartphone products, the smartphone manufacturers' own models, system scheduling, and app strategies must also be considered. It is difficult to replicate all results using a single prefetching strategy.
The good news, however, is that the hardware side is almost ready. In the exclusive interview, Huang Qiang stated that Longsys's UFS 4.1 standard products equipped with Huiyiwei's self-developed controller have achieved mass shipment and are ready for large-scale mass production and commercial use. The next step is to promote the integration of iSA into specific smartphone projects for further verification regarding power consumption, stability, and reliability.
Therefore, the main gap between iSA and its actual appearance in mass-produced smartphones is the final adaptation and engineering work.
For Longsys, this final step may be even more important than the previous ones. The 30B MoE on the reference design proves that this solution is entirely feasible, but what smartphone manufacturers really care about is how much DRAM can be saved after using iSA, how much the first-token latency can be reduced, and how much power consumption will increase during long-term operation. These data will only be truly convincing after running on mass-produced devices.
As for ordinary users, people do not care whether the expert parameters are placed in DRAM or NAND, nor do they care about how many times intelligent prefetching is performed behind this algorithm. What they can ultimately feel is simply whether the agent on the smartphone responds faster, whether the power consumption of using an AI smartphone for a long time will increase, and whether running AI in the background will affect the use of other apps.
If iSA can ultimately handle these issues well, the 30B MoE showcased by Longsys at this Snapdragon Summit might not just be a demo.
At least the next time we discuss on-device AI for smartphones, in addition to looking at the SoC and models, we might need to take an extra look at storage.