EN / 中文

10,000 Sandboxes in 10 Minutes: Intel Redefines Data Centers Built for Autonomous AI Agents

by bandaotichanyezongheng·September 29, 2026

By Fang Yuan

"10,000 sandboxes spun up within 10 minutes."

During Intel Connection 2026, Intel put forward this figure. The creation time for each sandbox must not exceed 60 milliseconds. A sandbox is not an empty shell; it must host a complete runtime environment, including images, file systems, and isolated spaces. This requirement is not a laboratory hypothesis; it comes from the production environments of real customers. It also points to a specific question: as agents evolve from chatting to actually getting work done, what exactly does a data center need to become?

Chen Baoli, Vice President of the Data Center Group and General Manager of China at Intel, made this judgment during the keynote speech: by the end of 2026, over 80% of enterprise applications will embed at least one AI agent; by 2031, the number of active agents in Chinese enterprise scenarios will reach 350 million. The daily token invocation volume for large models in China has increased to 5,000 times that of early 2024, and it continues to grow. Behind these numbers lies a profound adjustment taking place in the computing power structure of data centers.

01. AI Enters Its Third Phase, and the CPU Returns to Center Stage

In an interview, Chen Baoli detailed the workflow of an agent: after receiving a task, an agent first plans the task, breaking down the objective into several steps, and then invokes tools for execution. The intermediate results generated during execution need to be saved for reference in subsequent steps. If the task is complex, multiple sub-agents will be launched to collaborate in parallel, each with its own execution environment and context to remember. He gave a daily example: a user asks the agent to collect and summarize the previous day's AI news every morning at 8:00 AM. The agent needs to search the internet for information, and the retrieved content could be web pages, PDF reports, or videos. Information in various formats must be extracted and processed first before being handed over to the large model for analysis. Much of this preprocessing work is well-suited for the CPU. Moreover, a single model invocation is often not enough; the agent must break down tasks, invoke tools, and then decide the next steps based on the results.

He summarized the evolution of software engineering driving agent development into four stages: from Prompt Engineering to Context Engineering, then to Harness Engineering, and finally to Loop Engineering. Each step reduces human intervention, allowing agents to complete complex tasks more autonomously. Within these software mechanisms, many tasks are well-suited for execution by the CPU. Chen Baoli said, "This is also why the rise of AI agents over the past year has brought about a massive demand for CPUs."

Market expectations have also been significantly re-evaluated accordingly. Estimates for the CPU market size in 2030 were revised upward from USD 59 billion in February to USD 210 billion in August within just half a year, approximately 3.6 times the original expectation. The proportion of AI workloads in data centers is expected to increase from 40% in 2025 to 80% in 2030. Lip-Bu Tan, CEO of Intel, mentioned during the first-quarter 2026 earnings call that the focus of AI workloads is shifting from training to inference, and the CPU-to-GPU ratio is evolving from the previous 1:8 to 1:1. At this Connection conference, this trend was further corroborated: some complex agent scenarios even require higher-density CPUs. Chen Baoli also revealed a real data point from a customer during the interview: a leading domestic large model vendor saw its CPU demand increase fivefold from last year to this year.

02. The "Scratchpad" and "Isolation Chamber" for Agents

The working memory of an agent can be understood as a scratchpad. During reasoning, coordination, planning, and dialogue, the agent needs to retain intermediate results, which is the KV Cache. The problem is that as context windows grow longer, the volume of this "scratchpad" is exploding. A paper by Kimi shows that in agent scenarios, the hit rate of the KV Cache can reach 50% to 70%. This means the caching mechanism is indeed effective, but it also implies immense storage pressure. Chen Baoli provided another set of figures: for a 235B parameter model, if given a context window of 1 million tokens, the generated KV Cache is about 300GB, which already exceeds the HBM capacity of many GPUs. This either cannot fit or is not cost-effective, so the KV Cache must be hierarchically offloaded from HBM VRAM to DDR memory or even SSDs.

Intel's solution is the KV Cache Intelligent Acceleration Suite, with its core component called KV Shrink. This solution consists of a three-layer structure. The first layer is data movement, achieving efficient copying between VRAM and system memory via the DSA accelerator, delivering performance nearly 10 times that of CUDA solutions. The second layer is compression; the QAT engine can achieve a compression ratio of over 50% for KV Cache data, and lossless compression can also reach 1.42 times. The original 300GB KV Cache can be compressed to about 200GB before being written to disk. The third layer is system optimization and scheduling, which parallelizes the decompression of each layer with the computation of the previous layer, hiding latency within overlapping pipelines. Customer validation data shows that the KV Cache Intelligent Acceleration Suite can boost the KV Cache transfer rate by up to 9.66 times and improve time-to-first-token performance by up to 15 times. It supports two deployment modes: one is to place it directly in the head node of an AI server, offloading the KV Cache from VRAM to system memory for compression or even storage, and decompressing it back to VRAM when needed; the other is to use RDMA to transfer the KV Cache from VRAM to a remote server.

Another technical component is the sandbox. The essence of a sandbox is a lightweight virtual machine, or MicroVM, which trims and customizes the OS image and file system, aiming for speed, lightness, and lower memory footprint. It needs to protect two things: ensuring the agent's own sensitive information is not visible to other workloads, and preventing the system from being affected by the unpredictable code executed by the agent. The core KPIs for sandboxes are high concurrency and low latency. Under the constraint of a 200-millisecond startup latency, a single Xeon 6+ processor can support the concurrent creation of up to 350 sandboxes, which is 1.46 times that of a certain comparable product. A single processor can support the stable deployment of over 1,000 agents at most, relying on CPU overcommitment, where multiple agents are deployed on a single core simultaneously, time-sharing system resources.

The performance of Xeon 6+ in these two scenarios points to the CPU's intensive execution capability. It does not downclock under heavy system loads; over 200 cores work simultaneously, delivering nearly linear performance. This is the foundation for high-density sandbox deployment. From a microarchitecture perspective, an agent's invocation behavior is intermittent: it initiates an invocation, waits for the result, and then initiates the next one. In this mode, data might remain in the cache, but instructions may have been swapped out due to the intermittency. Intel adopts a strategy of multi-core sharing of large-capacity L3/L2 in its cache design; the L2 cache is 4MB shared by 4 cores, with up to 4MB per core, and 80% of instructions hit in the L2 cache.

03. Architectural Choices for Xeon 6+

In response to these demands, Intel's technical response centers on the Xeon 6+. The Xeon 6+ is the first data center CPU based on the Intel 18A process, featuring up to 288 E-cores, providing up to 8000 MT/s DDR5 memory bandwidth and 576MB of last-level cache. Intel 18A has entered mass production, and the high-performance version, 18A-P, has also advanced to the risk production phase.

However, what is even more noteworthy is the architectural choice. There is another approach in the industry: using a general-purpose CPU with external accelerator cards, or having the GPU take on more of the CPU's work. Intel chose to integrate acceleration engines such as QAT (data protection and compression acceleration technology), IAA (in-memory analytics acceleration), and AMX (AI matrix multiplication acceleration) entirely within the CPU. When asked why Intel chose this path, Chen Baoli explained: "If the acceleration capabilities are externalized, the round-trip overhead of data moving between the CPU and the accelerator card would offset the cost savings achieved by the offloading itself."

In the KV Cache scenario, this logic is very clear. QAT is a built-in compression engine that can complete the compression and decompression of KV Cache data without occupying GPU resources. DSA acts like a built-in DMA controller, responsible for efficient copying between VRAM and system memory. AMX is used for AI data preprocessing and vector database acceleration. The core design of this solution is pipeline overlapping: the decompression of each layer is parallelized with the computation of the previous layer, hiding the latency of compression and decompression behind the computation. During the keynote speech, Chen Baoli also demonstrated the effect of AMX in data preprocessing scenarios. He took the example of AI collecting and summarizing financial news: the agent needs to extract text from materials in various formats such as web pages, PDFs, and videos before feeding them into the large model for analysis. With acceleration capabilities like AMX, vectorization and reranking performance reach 2 times and 4 times that of the baseline, respectively. On the GPU front, Intel launched the new-generation GPU Crescent Island, based on the all-new Xe3P architecture, designed for AI inference and Agentic AI scenarios. One of its features is large VRAM; through LPDDR5x technology, it can be equipped with large-capacity memory of 128GB, 256GB, or even 480GB for AI inference and KV Cache storage.

04. From a Single Chip to an Entire Rack

Chen Baoli summarized the future data center into three types of coordinated clusters: CPU servers or CPU clusters are responsible for agent execution and work orchestration, GPU servers handle model inference, and storage clusters accommodate the ever-increasing data such as long conversations, documents, and contracts. "These three types of clusters share a very important commonality: they all require the CPU for data scheduling. AI factories are no longer measured solely by token output, but are shifting towards being task-oriented, achieving lower resource consumption and higher operational efficiency through system-level collaborative optimization."

Why is a full rack needed? Because when the number of agents expands from one hundred to one thousand, the bottleneck is no longer the core count of a single processor, but the system-level balance of memory capacity, inter-node communication, power supply, and cooling. The limit of sandbox overcommitment lies not in the CPU core count, but in memory capacity. Calculating with a 128-core system, a 1:2 core-to-memory ratio, and 4x overcommitment, a single node requires 1TB of memory. In the interview, Chen Baoli further explained the memory challenges brought by agents. In the traditional model, a virtual machine has a memory capacity of 2GB, with several cores corresponding to one VM, and each CPU core uses only a few hundred megabytes. Now it is the reverse: one core has to host four or five agents, and each agent requires 2GB. "Although it is not running every single moment, it consumes your storage capacity every single moment." This requires that within the limited CPU edges, two DIMMs can be plugged into each channel instead of just one, and even higher bandwidth must be maintained when plugging in two simultaneously.

We also saw that Intel, together with industry partners at the conference, released the Open Agent Rack Platform Specification, enabling more vendors to build their own rack products based on this standard. The framework for measuring the number of deliverable agents in a full rack is: the theoretical value multiplied by three coefficients, namely CPU scheduling effectiveness, resource balance, and compute offloading benefits. In terms of resource balance, tests found that the memory bandwidth requirement for each agent does not exceed 1GB/s, and it is recommended to configure at least 2GB of memory per agent. Compared to I/O, memory is more likely to become the bottleneck; and between memory capacity and bandwidth, capacity is more worthy of attention. Intel Xeon 6 and Xeon 6+ provide a CXL-based Flat Memory Mode, expanding memory via CXL. The hardware manages data placement between expanded memory and local memory, which is transparent to the operating system and applications, requiring no changes to the applications.

05. The Compute Ledger Is Being Rewritten

Placing the above technical details back into the industrial context, a more macro-level change is taking place.

The standards for measuring AI infrastructure are changing. In the past, the focus was on how many tokens could be generated; now, it is on how many effective tasks can be stably completed. This shift pushes the CPU to the position of the system hub. Task decomposition, tool invocation, sandbox scheduling, context management, and data preprocessing—these are tasks that GPUs cannot and should not do. What Intel is doing at the Connection conference is laying the hardware foundation for this shift. The fully built-in acceleration engines of Xeon 6+ ensure that data movement and compression remain within the CPU's scheduling scope; the hierarchical offloading of KV Shrink allows the limited HBM VRAM to carry more concurrent tasks; and the Open Agent Rack Platform Specification aims to expand single-chip capabilities into system-level capabilities.

There is another easily overlooked support point for this foundation: the ecosystem. Chen Baoli cited a case of his own: he installed a desktop agent and asked it to remove the background from a photo. The agent found that there was no background removal software on the computer, so it went online to find open-source PyTorch code, downloaded it, and wrote a background removal program by itself. "On the path of AI learning, all execution is done by the CPU," he added. That PyTorch code might have been written by someone over the past three to five years, and it is highly likely based on the x86 architecture. "Over the past 10 or 20 years, many features in the open-source community have been written based on x86. Therefore, in the agent era, using x86 to build agents has a significant advantage."

Returning to the question at the beginning: as agents evolve from chatting to actually getting work done, what exactly does a data center need to become? Intel's answer is that it needs to transform from a computing facility measured by token output into a system reorganized around the task rhythm of agents. The positioning of chips is changing, the role of memory is changing, the design logic of racks is changing, and the CPU is returning to the scheduling hub of this system. At the end of the interview, Chen Baoli concluded: "Over the next five years, the positioning of every chip in the data center will be re-discussed." The core issue of this discussion is not who has stronger computing power, but who is more suited for the division of tasks in the agent era. Judging from current industry trends, Intel is providing its answer in its own way.