On September 22, the first Immersive Consumption Carnival and the 2026 AI Terminal Innovation and Development Conference officially opened in the Beijing Economic-Technological Development Area (Beijing E-Town). Ji Chaohui, Vice President of Marketing for AMD Greater China, attended the conference and delivered a keynote report. Ji Chaohui pointed out that when inference becomes the primary workload for AI, token costs become an "unbearable burden" for enterprises. However, if agents are deployed on local new computing devices such as Agentic PCs and Agent Computers, leaving only heavy tasks to cloud-based large models, token costs can be reduced by 80%. "This is just like most of the time, we buy ingredients and cook at home, and only eat out on important occasions."
Agents Lead the Industry into the "Second Half" of AI, with Token Consumption Becoming an "Unbearable Burden"
Over the nearly four years since the launch of ChatGPT on November 30, 2022, the development of the AI industry can be divided into two stages. In the first three years, global large model companies have been striving to make AI smarter, enabling it to serve users in production, work, and learning, driving the AI training computing volume to expand continuously at a rate of 5 times per year.
In the past year, "putting AI to use" has become the focus of the industry. Therefore, while the computing volume for large model training continues to grow, the workload for AI inference has also risen rapidly. In 2025, the proportion of AI inference and AI training workloads reached 5:5, and in 2026, the proportion of AI inference has reached 60%. By 2030, the proportion of AI inference will continue to increase, reaching over 80%.
"We have entered the second half of AI, shifting from making large models smarter and smarter to the global effort to put AI to use," Ji Chaohui stated.
Agents are the key to AI's transition from conversational robots to productivity tools. They are not only the core paradigm for "putting AI to use" but also the turning point for AI to enter the fast track across thousands of industries. The working mode of ordinary large models is simple: users input questions, and they provide answers. Agents, however, are much more versatile. When given a task by a user, they use reasoning, sub-agent invocation, data processing, task dispatch, and tool invocation to ultimately feedback an execution result to the user.
However, the autonomy of agents comes at the cost of a sharp increase in token consumption. Gartner data shows that agent applications consume 5 to 30 times more tokens compared to conversational robots. IDC data indicates that by 2027, the AI token consumption of the Global 2000 enterprises (G2000 enterprises) will increase by 1,000 times compared to current levels. Under such a trend, tokens are increasingly becoming a cost challenge for enterprises.
Using Agents Like Cooking at Home, with Edge AI Taking the Lion's Share
Regarding the phenomenon that enterprises find it hard to afford token costs, Ji Chaohui made a vivid analogy: "The way enterprises use AI today is like eating at a restaurant for every meal—most tokens for AI applications come from cloud token suppliers, and we need to buy tokens and APIs from them. But when we live our daily lives at home, do we really eat out for every meal? Actually, most of the time, we buy ingredients and cook at home, and only eat out on important occasions."
So, how can we use agents like "cooking at home"? Relevant research shows that if daily auxiliary office work is placed on small models on AI PCs, and the agent applications required by agent AI employees are placed on medium models on Agent Computers, using only large inference models in the cloud for complex reasoning tasks, enterprise AI token consumption can be reduced by over 80%—using local models on AI PCs and Agent Computers is equivalent to "cooking at home" normally, while using cloud large models is "going to a restaurant."
Implementing this method of selecting models and computing locations based on tasks is inseparable from edge AI. It not only solves the problem of increasingly expensive token consumption but also meets users' needs for compliance, security, privacy, IP protection, immediacy, and data sovereignty. Most importantly, the capabilities of edge models have significantly improved over the past few years, making edge AI possible.
"We can see that the capabilities of the latest locally deployable Qwen3.8 model (with 27B parameters) have reached 80% to 90% of the most cutting-edge cloud large models, which is sufficient to drive the use of edge agents. The arrival of agents is triggering a new wave of evolution in AI terminals," Ji Chaohui said.
Launching New AI Computing Devices and Selecting AI Use Case Leaders
In March this year, AMD introduced Agent Computer as a new computing category, followed by the Agentic PC. In this speech, Ji Chaohui further elaborated on the usage of the two devices: the Agent Computer is for agents to use, while the Agentic PC is for agents and humans to use together.
Compared to traditional PCs that require continuous operation by humans via keyboards and mice, the Agent Computer only requires users to dispatch tasks (for example, sending them via Feishu) to work uninterruptedly for 24 hours until the task is delivered.
"The Agent Computer places agents on the edge rather than in the cloud. Future enterprises might have 1,000 human employees working alongside 10,000 agent AI employees. Some AI employees will run in the cloud, but due to compliance, privacy, data security protection, and data sovereignty, most of them are required to be deployed on the edge," Ji Chaohui said.
To build an Agent Computer, the first step is to "load" a foundational large model (at least 27B) and its required 24GB or more of VRAM to support the operational needs of the agent. Having such model capability is equivalent to hiring a Ph.D. graduate, but lacking work experience is not acceptable; therefore, a private knowledge base containing 30 years of experience in the relevant industry is also needed, along with over 20GB of storage space required for RAG KV cache expansion. Next, AI employees such as finance, customer service, human resources, and engineers need to be deployed. Each AI employee requires more than 10GB of VRAM, and they must be equipped with a strict SOP (Standard Operating Procedure)—only then can such a device meet the framework of an Agent Computer.
"If you use an Agentic PC, its AI capabilities can double your work efficiency. If an enterprise deploys an Agent Computer, it can increase the team's work efficiency by 10 times. If an enterprise adopts an agent workstation or server, it can boost the efficiency of the entire enterprise by 100 times. In the future, agents will play a crucial role in enhancing enterprise capabilities," Ji Chaohui stated.
Although the trend that AI will change future work models is clear, many users still lack a concrete concept of how to use AI to "multiply" their capabilities. To this end, AMD organized the Ryzen AI Agent Innovation Application Competition, aiming to screen out "pioneers" who have achieved 3x, 5x, or even 10x productivity through AI as examples. For instance, Cao Dongdong, a metal processing salesperson, found it difficult to provide accurate quotes when taking orders on e-commerce platforms because the processes and costs of various CNC machine tools in the backend factories varied. The Agent Computer handed over all the process flows, working hours, and cost information of all machine tools in the backend factory to the agent. As long as the customer's drawings were given to the Agent Computer, a clear and complete quote could be obtained. High school student Sophia Liu used the Agent Computer to create agent-assisted teaching, helping teachers implement layered teaching and teach students according to their aptitude.
"Each of us will become a super individual in the AI era," Ji Chaohui said.
Author | Zhang Xinyi
Editor | Wu Lilin
Art Designer | Ma Liya Supervisor | Lian Xiaodong