EN / 中文

213:1 Token Imbalance: Agent Workloads Redefine China’s AI Computing Infrastructure at AICC2026

by xindongxi·September 28, 2026

Author | Jiang Yu, Editor | Mo Ying

On the afternoon of September 21, at the Zhongguancun International Innovation Center in Beijing, Wang Shuwen, a core contributor to the open-source inference framework SGLang, released a set of data: in Agent traffic, the median input for a single request is 88,000 Tokens, while the output is only 413 Tokens, with the two differing by 213 times. "In the era of agents, computing power is not spent on output, but on the historical context carried repeatedly," he said. This took place at the "AI Software Stack" sub-forum under the AICC2026 AI Computing Conference. At the scene, the SGLang team shared insights on the changes in inference workloads in the Agent era.

On the same day, the AICC2026 AI Computing Conference was held at the Zhongguancun International Innovation Center in Beijing, attracting nearly 3,000 industry professionals to the scene. The conference set up seven thematic forums in one go: AI Chips, Token Factory, Open Supernode, Large Models, AI Software Stack, Embodied AI and Intelligent Computing Development, and AI Work Agents, bringing together different segments of the industry chain including chips, systems, software, models, and applications.

On that very same day and at the same conference, chip manufacturers, system vendors, and inference framework developers all gathered around the same table. From the cache pressure brought by long contexts, to the adaptation of diverse chips, and then to the collaboration of interconnects, scheduling, and software stacks, the focus of the discussion can hardly be broken down into a problem of a single segment. They ultimately point to the same question—as agents move from a single invocation to continuous task chains, what should be the next step for China's AI computing?

01.The Explosion of Agents Every Layer of AI Computing Is Being Redefined

The explosion of agents is pushing computing power to the position of national strategic infrastructure. In 2026, agents have moved from demos at launch events into real-world production, scientific research, and social services; a single agent task may involve planning, multi-turn reasoning, tool invocation, memory, verification, and retries, resulting in longer workloads and increased concurrency, while the system must also be recoverable, observable, and cost-effective. At the AICC scene, the conference broke down this change into three layers of industry signals for discussion.

First, agents have changed the service targets of computing.

Infrastructure used to serve a single model invocation, and now it must serve a continuous task.

Second, real-world application workloads are inversely defining the computing architecture.

In June 2026, the Jalapeño chip released by OpenAI and Broadcom directly guided chip design with model routes, computing kernels, inference services, and product requirements. The demands of models and applications are extending upstream along the industry chain.

Third, the diversification of technology routes and the deepening of system-level integration are occurring simultaneously on a global scale.

At Hot Chips 2026, GPUs, custom chips, wafer-scale computing, RISC-V, Chiplets, and processing-in-memory appeared on the same stage. Meanwhile, leading companies are strengthening full-stack collaboration from chips to software around real workloads. In the context of China's AI computing industry, AI is extending in two directions simultaneously: one upwards, where inference, tool usage, long-horizon tasks, and scientific discovery capabilities continuously break through the upper limits of intelligence; the other outwards, where cost reductions and the popularization of agents are transforming intelligence from a capability held by a few institutions into a productivity universally needed by enterprises, individuals, and society. What agents are changing is the organizational model of the entire AI computing industry chain. Understanding this is the starting point for comprehending this conference and the infra layer focused on below.

02.From Answering a Single Question to Continuous Execution: What Happens on the Inference Side During an Agent Task

Returning to the set of data from Wang Shuwen at the beginning, this directly reveals a change: Agents are turning inference from a single invocation into a continuous task chain. In the past, a single inference was just "answering a question": a request comes in, the model computes once, Tokens are returned, and it ends. Now, agents need to "work continuously": to complete a task, they invoke the model over and over again, planning first, then searching, then invoking tools, then reading results. The context accumulates longer and longer, and every time the model is invoked, the complete historical context must be input. What the inference system has to face has changed from independent computations one by one to a continuous task chain.

The form of the workload has changed, and the computing pressure has changed accordingly. The 213:1 input-output ratio means that in agent scenarios, a large amount of historical information will repeatedly enter the model with each round of invocation, while the truly newly added content actually accounts for only a very small proportion. These workloads are highly concentrated in three patterns: multi-turn reuse and appending, sub-agent forking, and tool looping and retrying. When these three patterns overlap, the data structure is "more like a tree": the user has only chatted with the agent for a few rounds, but the agent backend has already interacted with the model dozens of times. In the end, over 90% of the Tokens are spent on the accumulation of historical context, and only a small segment truly needs to be recomputed.

According to a report by SemiAnalysis, agents now already contribute about 70% of inference traffic. As the workload changes, the optimization goals of inference systems must also change accordingly. This is exactly what open-source inference frameworks like SGLang have been doing for years: since a large number of Tokens are repeatedly computing the same context, the solution is to find a way to "reuse" them rather than recompute.

The RadixAttention, first proposed by SGLang, uses a prefix tree to look up the KV Cache, allowing Requests with shared prefixes to quickly find the already computed KV Cache through the prefix tree and reuse it; when the VRAM cannot hold it all, HiCache is used to offload the cache to main memory or even SSDs, extending the context even further. The effect of this change is tangible. The field test data presented by Wang Shuwen on-site showed that the early cache hit rate was only 10% to 30%, but after integrating cross-instance storage, it can be increased to 80% to 90%, and the time-to-first-token latency also decreased significantly. From answering a single question to continuous execution, the core contradiction on the inference side has shifted from computing fast to computing with less repetition. Only by understanding this can we see what the next two layers—the inference framework and the operator system—are respectively solving.

03.From Model Release to Running in Production The Infra Layer Must Pass "Nine Checkpoints"

What users see is just a "usable" model, but from its release to truly running in the industry, there is a long pipeline in between. Applications and agents propose requirements, models are released, inference frameworks provide support, operators are completed, chips are adapted, functions are run through (Day 0), accuracy is aligned, single-operator/single-machine performance is tuned, system-level inference optimization is performed, and finally, it goes into mass production. At the AICC scene, this chain was repeatedly sorted out by multiple speakers, and the industry calls it the "nine checkpoints".

This pipeline can roughly be divided into two layers. The upper layer is the inference framework, answering how quickly inference services can be up and running when a new model is released today; the lower layer is operators and systems, answering what needs to be solved for a model to run on domestic diverse chips with correct results and good enough performance. These two layers each had a visible milestone progress before and after the conference. On the framework side, the answer lies in "Day-0"—the inference framework can complete support on the very day the new model is released. A core contributor to SGLang listed on-site the models that were supported at the first opportunity recently, including DeepSeek-V4.1, Qwen3.8 Flash, and Nemotron 3.5. On the operator side, the change happened less than a month before the conference.

Alibaba's Qwen3.8-Flash-Next, Zhipu's GLM-5.3-Flash, and Tencent's Hy4 preview, three MoE models, were launched one after another, bringing new adaptation challenges such as brand-new attention structures, complex hybrid attention deployment, and ultra-large parameters of 770B. The Zhizhi FlagOS community, led by BAAI, completed three rounds of multi-chip adaptation within four days, two of which achieved Day-0 on the day of release. The adaptation scope covers 10 AI chips including Moore Threads, Muxi, Kunlunxin, Hygon, Iluvatar CoreX, and Tsingmicro.

Behind this speed is a set of gradually accumulated public technical assets, including operator libraries, compilers, and a plugin layer that interfaces with vLLM and SGLang. After a new model emerges, common capabilities can be accumulated in the public layer, while differences between chips are converged as much as possible into the plugin and underlying adaptation links, thereby reducing the cost for each party to repeatedly complete the same set of work. We observed on-site that this pipeline, which was previously fought by various manufacturers independently, is turning into a set of reusable public infrastructure.

04.Rising Demand and Constrained Supply The Answer for China's AI Computing Is Open-Source Collaboration

Of course, for the Chinese market, this pipeline has another more realistic constraint. As agents accelerate their entry into production, scientific research, and social services, inference demand continues to grow, posing higher requirements for the scale, efficiency, and adaptation capabilities of computing power supply. How to solve this problem? A common direction we saw at the conference is open-source and diversification. Around this direction, the industry is forming multiple exploration paths, among which open-source frameworks, basic software, and diverse computing power adaptation is a route worth paying attention to: organizing different chips, systems, and software into stable, efficient, and scalable system capabilities. Specifically, two types of samples can be observed. Open-source inference frameworks such as SGLang are a microcosm of one of them. Through caching, scheduling, and system-level optimization, inference frameworks are further improving the utilization efficiency of computing resources, allowing limited hardware resources to carry more inference tasks.

The other type of exploration occurs at the diverse chip adaptation layer. Taking FlagOS led by BAAI as an example, it hopes to reduce the adaptation costs between different models, frameworks, and AI chips through a more unified and open basic software system, enabling new models to run faster on different computing platforms. The FlagOS community summarizes this problem as evolving the complex "M×N" adaptation towards an "M+N" model as much as possible. In the past, each model framework and each chip might require separate adaptation; if public layers such as operators, compilation, and frameworks can form more unified interfaces and collaboration methods, and the framework side and hardware side interface with the public layer respectively, there is an opportunity to reduce repeated adaptation and improve the efficiency of technology reuse.

The value of this open-source collaboration is also beginning to be reflected in specific numbers. We heard a vivid metaphor on-site, which views inference services as a "Token factory": the inputs are fixed costs such as GPUs, servers, networks, and electricity, and the output is the variable revenue of Tokens. Some manufacturers have built an inference foundation based on open-source inference frameworks and distributed caching systems, with daily average Token production stably exceeding 1 trillion, production efficiency increased by 3 times compared to before, and total capacity growing by more than 30 times. It can be seen that under the situation of continuously growing computing power demand and constrained hardware supply, China's AI computing is attempting to use open-source collaboration to organize limited single-card performance, scattered diverse chips, and rapidly changing models into more stable, more efficient, and scalable inference capabilities.

05.Conclusion: Making Intelligence Accessible and Affordable Is the Next Proposition for China's AI Computing

In the era of agents, the true proposition for China's AI computing is how to organize diverse chips, systems, and software to continuously support the large-scale operation of agents—both allowing frontier intelligence to continuously break through upper limits and enabling intelligence to universally enter thousands of industries. The answer given by AICC2026 is openness and collaboration. At the conference scene, the progress presented at the inference framework and operator layers may seem technical and subtle, but they are actually all answering the same question: how to make a new model run faster, run on more chips, and be used by more people at a lower cost.

When agents transform from demos by a few institutions into a production capability universally needed by enterprises, individuals, and society, the competition in AI computing has escalated into a competition of open collaboration. Whoever can twist diverse chips and software into a scalable supply of computing power will be able to define the foundation for universal intelligence. This is exactly the next step being written by this conference and China's AI computing industry.