In September 2018, at the Apsara Conference in Hangzhou, Alibaba announced the establishment of a chip company named T-Head. The name is said to originate from an African trip taken by the man himself.
To be honest, not many people in the industry were optimistic at the time. An internet company making chips sounded a bit like a cross-over hobby project. The chip industry is notoriously slow, requiring long development cycles and heavy capital investment, while also demanding acceptance of a high failure rate: it usually takes several years from project initiation to mass production for a single chip, and the cost of a single tape-out can be tens of millions or even over CNY 100 million. Once a mistake is made, neither the money nor the time can be recovered.
Many believed that an internet company accustomed to the "small steps, fast running, and rapid iteration" approach might not have the patience to stick with it.
Eight years later, back at the Apsara Conference, the same company, T-Head, released the new generation of training and inference integrated AI chip, Zhenwu V900. Its performance is three times that of the Zhenwu M890, making it one of the most powerful AI chips independently developed in China today, capable of meeting the training and inference needs of trillion-parameter large models.
In recent years, "computing power" has become the most anxiety-inducing term in the entire AI industry. As models grow larger and GPUs become harder to buy, every company is asking the same question: where exactly will our computing power come from? Looking back at these eight years in this context, T-Head, which was once questioned as just playing around, has quietly developed almost all the key chips in the data center.
Eight Years: From Single-Point Breakthroughs to Full-Stack Independent Development
Let's first take a look at T-Head's eight-year chip development path.
In 2019, T-Head, just one year after its establishment, launched its first chip, Hanguang 800. This is an inference chip deeply customized for specific scenarios, capable of processing 78,000 images per second. At the time, it ranked first among similar chips in both performance and energy efficiency ratio, and was later used in Taobao's main search during the Double 11 shopping festival.
In 2021, the Yitian 710 server CPU was released, making Alibaba the first internet company in China with CPU R&D and design capabilities. This chip's performance is 20% higher than the industry benchmark at the time, and its energy efficiency ratio is improved by over 50%. It is still running workloads such as video transcoding, high-performance computing, and gaming on Alibaba Cloud.
In 2023, the first SSD controller chip, Zhenyue 510, was added on the storage side.
In 2024, the first training and inference integrated AI chip, Zhenwu 810E, was launched. Paired with the independently developed T-Head SAIL software stack, it supported the training and inference of the Qwen large model.
In April this year, the first smart NIC, Panmai 920, was released on the network side. This is the first 400G smart NIC in China with a built-in PCIe Switch.
In May this year, the Zhenwu M890 was unveiled, with performance three times that of the 810E, and a 128-card super node equipped with the ICN Switch 1.0 was also introduced.
In September this year, the Zhenwu V900 was released, with performance again three times that of the M890.
AI chips, server CPUs, interconnect chips, smart NICs, and storage controllers—in eight years, computing power, networking power, and storage power have all been covered. The Zhenwu series has iterated through three generations in three years and has served over 650 enterprise customers. At the Apsara Conference, Eddie Wu, CEO of Alibaba Group, stated that due to the maturation of T-Head's chip product line and its widespread application by customers, T-Head's annual shipment volume of AI chips will significantly increase.
From Making Components to Building Systems
Looking back at this timeline, this year marks a clear turning point.
In the past few years, what T-Head did was more like making components: filling in whatever was missing, developing the key chips in the data center from scratch one by one.
This year, these components have begun to be assembled into a complete machine.
Why now? Because AI's demand for computing power has changed. Large models have reached the stage of trillion parameters, MoE architectures, and the popularization of Agent applications. Simply comparing the specs of a single chip is no longer enough; what truly matters is the efficiency of the entire computing system.
If an AI computing system is broken down, it roughly has three layers: the bottom layer is the computing power of a single chip, the middle layer is the interconnection capability between chips, and the top layer is the general-purpose computing power responsible for scheduling and execution.
The three chip-related advancements at this Apsara Conference correspond exactly to these three layers: Zhenwu V900 solves "computing fast," ICN Switch and super nodes solve "connecting well," and the Yitian CPU roadmap solves "scheduling effectively."
Let's look at them layer by layer below.
Zhenwu V900: How to Have It Both Ways in Training and Inference
The Zhenwu V900 is a training and inference integrated chip. Many current AI chips have a division of labor, focusing either on training or inference, because the hardware requirements for these two tasks differ significantly.
Training is a memory-intensive process. Taking the common mixed precision plus Adam optimizer as an example, each parameter, besides the weights themselves, needs to store gradients, FP32 master weights, and two optimizer states. Roughly calculated, each parameter takes up about 16 bytes. For a trillion-parameter model, just these states require about 16TB, not to mention the intermediate activation values. Therefore, training must partition the model across a large number of chips. The larger the memory of each chip, the fewer partitions are needed, and the less communication between chips is required.
Inference, on the other hand, is a bandwidth-intensive process. For every token generated by a large model, the relevant weights must be read from memory, while simultaneously maintaining the ever-growing KV Cache. The bottleneck at this stage is often not computation, but memory capacity and bandwidth. Therefore, inference tries every way possible to reduce precision, from FP16 to FP8 and then to FP4. Every time the precision is halved, the number of parameters that can fit in the same memory doubles, and the amount of data read is also halved. Of course, reducing precision to FP4 can easily degrade model performance, so the industry generally pairs it with techniques like block-wise scaling to maintain accuracy.
Training and inference integration means having a single chip simultaneously meet these two almost contradictory sets of requirements. Based on T-Head's independently developed parallel computing architecture, the Zhenwu V900 is equipped with 216GB of large-capacity memory, and the inter-chip interconnect bandwidth reaches 1200GB/s, which is in the same order of magnitude as NVIDIA's latest Blackwell Ultra (288GB memory, 1.8TB/s NVLink bandwidth). It natively supports the full spectrum of scenarios from high-precision training to low-precision and ultra-low-precision inference. Compared to the previous generation M890, the single-chip performance of the V900 is improved to 3 times.
The M890 was just released in May this year, and four months later, performance has tripled again, indicating that T-Head is not just making incremental updates.
According to Alibaba, the V900 will be mass-produced and sold in the first quarter of 2027.
The Real Bottleneck: The Communication Wall
However, no matter how strong the single-chip performance is, it only makes the longest plank of the wooden barrel longer. The short plank that now determines the performance of AI systems is the communication capability between chips, which is the legendary "communication wall."
This bottleneck did not emerge overnight but is closely related to changes in models and chip connection methods.
In traditional designs, an AI server typically houses 8 AI chips. Within the same server, chips are directly connected via dedicated high-speed interconnects, with bandwidth reaching hundreds of GB/s or even over 1TB/s. However, once stepping outside this server, chips can only communicate over the network via NICs and switches, with bandwidth at only about tens of GB/s, which is more than an order of magnitude lower than intra-machine interconnects, and latency is much higher.
To use an analogy, chips in the same server are like colleagues sitting in the same office; they can just turn around and talk when there is an issue. Chips across servers are like working in different buildings; every communication requires a phone call or an email. There is industry jargon for this: the former is called Scale-up, and the latter is called Scale-out.
The problem is that the most communication-intensive operations in large models, such as tensor parallelism and the expert parallelism mentioned earlier, require frequent data exchange between chips. Using the previous office analogy, these experts can only run efficiently if they are in the same office. When the model scale was not yet large, a small office composed of 8 chips was basically sufficient. But for trillion-parameter MoE models, experts often need to be distributed across dozens or even hundreds of chips. 8 cards are obviously not enough to hold them, so a large amount of communication has to go over the much slower network.
This is how the communication wall is formed.
The solution is actually quite straightforward: expand this office. The commonly used method now is to use dedicated interconnect chips to extend the scope of high-speed interconnects from 8 chips in one server to dozens, hundreds, or even thousands of chips, allowing them all to communicate with near-intra-machine bandwidth and latency. Such a system composed of a large number of AI chips, all internally connected at high speed, is the prototype of a super node.
Therefore, how many chips a super node can connect and how fast it can connect them largely determine the size of the model it can run efficiently.
Super Nodes: Making a Thousand Chips Act Like One
So, how exactly can thousands of chips be connected into a super node? This is the second highlight of this event.
The core of the super node is an interconnect chip dedicated to data exchange between chips, the ICN Switch, which stands for Inter-Chip Network.
This chip is not appearing for the first time. At the Alibaba Cloud Summit in May this year, T-Head released the ICN Switch 1.0 alongside the Zhenwu M890. It supports T-Head's independently developed ICN interconnect bus protocol, with a throughput reaching 25.6Tbps and point-to-point latency between chips below 150 nanoseconds. Alibaba used it to pack 128 M890s into a single cabinet, forming a 128-card super node, increasing the scale of full-bandwidth interconnects from 16 cards in the previous generation to 64 cards.
At this Apsara Conference, T-Head combined the Zhenwu V900 and the ICN Switch, pushing the scale of full-bandwidth high-speed interconnects to the thousand-card level, and the same architecture can cover all scenarios from a single machine with 8 cards to a thousand-card cluster. From 16 cards to 64 cards, and then to a thousand cards, it is equivalent to expanding the aforementioned office by dozens of times in just over a year.
Scaling up is only one aspect; more crucial is how the chips in this super node communicate with each other. It follows two very important principles: native memory semantics and unified memory addressing.
Let's talk about memory semantics first. Traditional inter-chip communication uses "message semantics": the sender needs to prepare data buffers, construct descriptors, and initiate DMA transfers. After receiving it, the receiver must perform completion notifications, synchronization, and then copy the data to the location needed for computation. Every step incurs software stack overhead, and a single communication often takes several microseconds. For large blocks of data, this overhead can be ignored; but for high-frequency small data packets like those in MoE, the time spent on these procedures might be longer than the actual data transmission time.
Memory semantics, on the other hand, allows one chip to directly access the memory of another chip using read and write instructions, just like accessing its own local memory, without any packaging, handshaking, or copying in between. In the M890 generation, the ICN Switch 1.0 had already reduced communication latency to the hundreds of nanoseconds level, which is an order of magnitude lower than traditional methods.
Next is unified addressing, which means that the memory of all cards is connected into a single piece in the address space, forming a globally shared large memory. To the software, it no longer sees a thousand mutually independent memories, but one giant memory.
To use an analogy, originally each workshop had its own warehouse. If you wanted to use parts from another workshop, you had to fill out forms, pack, ship, and sign for them. Now all warehouses are connected, shelves are uniformly numbered, and you just go directly to the corresponding shelf to get the part you need.
The biggest value brought by this feature actually lies on the software side. Those who have done distributed training should know clearly that deciding which card to place data on, when to move it, and how to overlap it with computation is very energy-consuming and error-prone. Unified addressing greatly simplifies the programming model, allowing framework and operator developers to focus more energy on the model itself. This is also the reason why global AI chip manufacturers are moving towards memory semantic interconnects.
From the 128-card super node of the M890 to the thousand-card full-bandwidth high-speed interconnect of the V900, the Scale-up domain has expanded by nearly an order of magnitude, and the same architecture can cover all scenarios from a single machine with 8 cards to a thousand-card cluster.
In other words, now thousands of Zhenwu V900s can finally work collaboratively like a single "super chip."
Computing, Storage, and Networking Assembled: A 500,000-Card Computer
Having computing power and interconnects is not enough. The new generation of super node servers released by Alibaba this time contains four types of independently developed chips: the Zhenwu V900 is responsible for computation, the ICN Switch handles Scale-up interconnects within the super node, the Panmai smart NIC handles Scale-out networking between super nodes, and the Zhenyue SSD controller handles data storage and retrieval.
Computing, storage, and networking are all completed at once, achieving system-level collaborative optimization.
Coupled with the newly designed intelligent computing center network architecture of Alibaba Cloud, this system can stably support a single AI cluster on the scale of 500,000 cards.
Additionally, super nodes based on the Zhenwu M890 have already been commercialized at scale, making it the first super node in China to run two models with over 2 trillion parameters: Qwen3.8 and Kimi K3. These two models you are currently using are very likely running on these T-Head chips. Through operator optimization and hardware-software collaboration, Zhenwu super node instances can achieve up to 1.5 times performance improvement in Agentic inference scenarios.
Without changing the hardware, gaining 50% more performance through system optimization—this is the value of a full stack.
CPU: The Underrated Protagonist in the Agent Era
The final layer is the CPU.
Over the past two years, when talking about AI computing power, people mainly focused on GPUs. But in the Agent era, the importance of CPUs is resurging. Task planning, state management, tool invocation, as well as the preprocessing and orchestration of massive amounts of data for agents, are almost entirely handled by CPUs.
The GPU is responsible for thinking, while the CPU is responsible for scheduling and execution. If the CPU cannot keep up, no matter how strong the GPU is, it can only wait idly.
Therefore, the CPU needed for Agents must have strong single-core performance, high memory bandwidth, and be tightly coupled with the GPU.
This time, T-Head publicly unveiled the roadmap for its Yitian server CPUs for the first time: launching the Yitian 720 and Yitian 730 in 2027. The 720 features comprehensive improvements in single-core performance, core density, and energy efficiency; the 730 adopts T-Head's fully independently developed CPU microarchitecture for the first time, with single-core SPECint 2017/GHz performance reaching up to 1.4 times that of the Yitian 710.
In the future, the Yitian 750 will continue to be launched, adopting the second-generation independently developed core and supporting T-Head's independently developed ICN inter-chip interconnect bus protocol. This is where it gets interesting.
Currently, most connections between CPUs and AI chips are via PCIe, which has relatively limited bandwidth and latency. In Agent scenarios, the CPU handles scheduling while the AI chips handle inference. Data interaction between the two is very frequent, making PCIe highly likely to become a bottleneck, and the communication wall reappears. Additionally, as context lengths grow longer, the industry is also trying to place part of the KV Cache in the CPU's large-capacity memory, which similarly requires a wider and faster channel between the CPU and AI chips.
The integration of the Yitian 750 into the ICN bus means that the CPU and the Zhenwu AI chips enter the same interconnect system. The super nodes mentioned earlier solved how to connect AI chips; this step solves how to connect the CPU and AI chips. At this point, the key large-chip puzzle piece in Alibaba's computing power landscape is also completed.
Define Requirements, Don't Guess Them
Returning to the question at the beginning: what exactly has T-Head, which was questioned as "just playing around" eight years ago, achieved now?
The answer given at this Apsara Conference is that they have built a computing system tailor-made for the AI era.
Zhenwu handles computing power, ICN Switch handles interconnects, Panmai handles networking, Zhenyue handles storage, and Yitian handles scheduling. In the future, CPUs and AI chips will also be connected to the same bus.
The competition for computing power has now shifted from competing on chips to competing on systems. Such a system must be repeatedly refined in real-world scenarios. Alibaba's biggest advantage in making chips is that it itself possesses massive and complete application scenarios: it has the cloud, Taobao, DingTalk, various actual businesses, and top-tier large models like Qwen. Once the chips are made, they can be directly deployed to run on Alibaba Cloud and its own businesses, verified and optimized using its own cutting-edge models, and then rapidly scaled.
Chips, models, platforms, and cloud—this complete chain is very difficult for the vast majority of chip companies to replicate.
From this perspective, many chip companies are guessing what everyone will need in a few years, while Alibaba is defining what it will need in a few years itself.
The doubts from eight years ago have received a quite complete answer at this year's Apsara Conference.
(Note: This article does not represent the views of the author's employer.)