I attended the Apsara Conference recently, and Alibaba Cloud's AI computing hardware left a deep impression on me.
Over the past few years, I have been closely following domestic computing power. I usually focus more on Huawei's Ascend and Kunpeng, as well as chips from vendors like Hygon, Moore Threads, MetaX, and Cambricon. To be honest, although I had heard of Alibaba Cloud's computing hardware, I hadn't paid much attention to it.
Alibaba Cloud's approach is quite different from the vendors mentioned above. As a cloud service provider, Alibaba Cloud acts as the client, much like telecom operators. Its hardware development is primarily for self-consumption, with most products not sold externally.
In other words, Alibaba Cloud uses its own hardware to build its own data centers (computing resource pools) and then sells (leases) computing power externally in the form of cloud services or platforms.
Take Huawei, for example. Although it has its own cloud, its primary role remains that of a vendor. It provides computing hardware and software solutions and products to cloud service providers, telecom operators, and government/enterprise clients. These clients then build their own computing clusters and data centers to run their own businesses or resell to downstream clients.
The advantage of Alibaba Cloud's approach is self-production and self-consumption with clear requirements: the cloud service department dictates what products it needs, and the R&D department designs accordingly, without needing to consider external demands or opinions. As long as the performance is strong enough and convenient for internal use, that is sufficient.
Since it is not sold externally, there is no need to worry too much about the ecosystem; providing upper-layer business services is enough.
However, the downside is that because it is not heavily oriented towards the open market, it cannot easily be sold to other clients, making it impossible to earn more "hardware revenue." From an ecosystem perspective, it is relatively closed, making it difficult for external developers to directly develop and adapt based on Alibaba Cloud's hardware.
If Huawei is somewhat like NVIDIA (selling shovels), then Alibaba Cloud is more like Google (making its own shovels and mining its own gold).
Looking at the exhibition at this Apsara Conference, I found that Alibaba Cloud is highly professional in hardware. Many hardware performance metrics have surpassed similar products from competitors. The richness of the hardware categories also exceeded my expectations.
Frankly speaking, it truly gives a feeling of seeing a completely transformed entity. Alibaba Cloud has been working quietly and suddenly produced so many self-developed products, which is indeed surprising.
Alibaba Cloud has too many types of AI basic hardware, and they particularly like giving them various strange "martial arts style" names, which always makes it hard to keep track.
Next, I will provide a detailed inventory of Alibaba Cloud's AI hardware product system. On the one hand, this will help everyone recognize the product names; on the other hand, it is a good opportunity to gain an in-depth understanding of the performance and characteristics of each product.
█ Chips
Alibaba Cloud's AI computing hardware infrastructure can be horizontally divided into three major categories: "computing power, storage power, and networking," and vertically into three levels: "chips, servers/super nodes, and full-rack clusters."
Let's first look at the chips.
Alibaba Cloud's chips are basically all based on T-HEAD.
T-HEAD, fully named T-HEAD Semiconductor Co., Ltd., is a wholly-owned chip design entity established by Alibaba Group in 2018, somewhat similar to Huawei's HiSilicon.
T-HEAD's self-developed chips form the foundation of Alibaba Cloud's computing hardware system, divided into two major product lines: data center computing chips and embedded chips (IoT, edge side).
● Zhenwu
Among the data center computing chips, the most important core chip is the Zhenwu series AI training and inference integrated chip, which can be benchmarked against Huawei's Ascend series.
Representative models of the Zhenwu series include Zhenwu 810E, Zhenwu M890, and Zhenwu V900.
Specific parameters are shown in the table below:
The most noteworthy is the Zhenwu V900.
Just released at this Apsara Conference, this chip is based on a self-developed parallel computing architecture, equipped with 216GB large-capacity VRAM, and features an inter-chip interconnect bandwidth of 1200 GB/s. It natively supports FP8 and FP4 low-precision computing, with performance 3 times that of the M890.
According to Alibaba Cloud's plan, the Zhenwu V900 will be mass-produced in Q1 2027.
The Zhenwu J900 is the planned next-generation flagship, with not much information available at the moment.
The predecessor of the Zhenwu series was the Hanguang series. The representative model is Hanguang 800, an early cloud inference chip targeting image, video, and OCR scenarios, which was deployed on a large scale.
● Yitian
Next is the Yitian series. Yitian is an Arm-based general-purpose CPU, similar to Huawei's Kunpeng series.
Representative models of Yitian include Yitian 710, Yitian 720, and Yitian 730. The 710 is already in commercial use at scale, while the 720 and 730 will be launched next year.
It is claimed that the Yitian 720 has comprehensively improved in single-core performance, core density, and energy efficiency. As for the Yitian 730, it will be the first to be designed based on T-HEAD's fully self-developed CPU microarchitecture, with single-core performance increasing up to 1.4 times that of the Yitian 710.
In the future, Alibaba Cloud will also launch the Yitian 750 based on the second-generation self-developed core.
If I remember correctly, the Zhenwu chips can be sold directly as bare chips externally. The Yitian seems unable to do so and is mainly for internal use.
● XuanTie
The XuanTie series targets edge and IoT (Internet of Things) applications and is based on the RISC-V architecture embedded CPU chip. This series is the founding product of Alibaba Cloud's T-HEAD, sold directly externally, with a very high shipment volume in China, seemingly the highest among RISC-V processors.
The XuanTie series includes three major product lines: C series, E series, and R series.
The C series is high-performance RISC-V, targeting high-performance computing for edge AI, intelligent cockpit, embodied AI, and smart vision. Representative models include XuanTie C908, C910, C920, C930 (main server-level), and C950 (flagship).
The E series is a low-power MCU (Microcontroller Unit), positioned for extreme simplicity, low power, and small area, targeting embedded scenarios.
The R series is positioned for hard real-time, low latency, and high determinism, targeting industrial control, automotive control, motor drive, etc.
● CIPU, Panmai, ICN
What was just mentioned are computing chips. T-HEAD also has communication, storage, and security chips.
For communication solutions, we need to talk about CIPU, Panmai, and ICN.
Alibaba Cloud proposed the CIPU concept a few years ago, which stands for Cloud Infrastructure Processing Unit, a self-developed cloud infrastructure processor by Alibaba Cloud. This chip replaces traditional NICs, offloading virtualization, networking, storage, and security to hardware, reducing the burden on the CPU and improving overall efficiency.
CIPU is actually somewhat similar to NVIDIA's DPU (Data Processing Unit) and Intel's IPU (Infrastructure Processing Unit).
Panmai is T-HEAD's self-developed smart NIC chip/NIC hardware, an ASIC chip for NICs responsible for high-speed transmission of massive data between GPUs, focusing on high-speed networking for AI computing clusters (Scale-out networking). It can be benchmarked against NVIDIA's ConnectX series smart NICs. The current flagship version is Panmai 920.
Alibaba Cloud also has a highly critical ICN Switch.
ICN, fully named Inter-Chip Network, refers to inter-chip interconnect switching (Scale-Up, intra-node interconnect, intra-rack interconnect), somewhat similar to Huawei's UnifiedBus technology and NVIDIA's NVLink technology.
ICN Switch 1.0 is used in the Panjiu AL128 super node (128 Zhenwu M890 cards per rack, introduced below), with a single-chip throughput of 25.6 Tbps, supporting 64-card full-bandwidth non-blocking interconnect.
The next generation of ICN Switch (2.0?), estimated to be launched in Q1 2027, paired with Zhenwu V900 and Panjiu AL64 super nodes, features an inter-chip bandwidth of 1200 GB/s, aiming to support 1,000-GPU scale full-bandwidth interconnect within the rack domain, with native memory semantics and unified addressing. I recall that Huawei's Ascend 950 super node (UnifiedBus 2.0) also supports these features.
● Zhenyue
For storage, T-HEAD has the Zhenyue series storage controller chips, with the representative model being Zhenyue 510 (SSD controller chip, PCIe 5.0, IO processing capability of 3,400K IOPS, latency of 4 μs, based on the RISC-V XuanTie core).
In terms of storage, there are also AliFlash and some SSD product hardware.
● Security
For security, T-HEAD has the Shendun Trusted Security Chip, as well as independent TPM/TCM Trusted Platform Module chips and hardware TEE encrypted computing.
█ Servers / Super Nodes
Next, let's look at the server and super node level.
Alibaba Cloud's self-developed cloud-native server brand is called Panjiu, divided into two main branches: general-purpose servers and AI super node servers.
General-purpose servers mainly target conventional business scenarios such as cloud computing and big data, including high-performance computing, high-performance storage, and large-capacity storage. No need to elaborate much.
The focus is on AI computing super nodes. Panjiu super nodes include AL64, AL128, and AL144, where the numbers represent the quantity of GPUs.
The latest Panjiu AL128 super node server is built-in with Zhenwu V900, next-generation ICN Switch, Panmai smart NIC, and Zhenyue SSD controller chip. It features an orthogonal cable-free design, Pb/s-scale Scale-up bandwidth, and sub-100-nanosecond latency. A single cluster can be scaled up to 500,000 GPUs and will go online in Q1 2027.
Panjiu AL144, Alibaba Cloud's flagship AI super node, focuses on 10,000-GPU scale Scale-Up optical interconnect. A single rack supports up to 144 GPUs (intra-rack via orthogonal copper interconnect, also with Pb/s-scale Scale-up bandwidth and sub-100-nanosecond latency). Cross-rack can be scaled to 10,000 GPUs (via secondary SNPO optical interconnect, scalable up to 10,368 GPUs), targeting MoE ultra-large models.
In terms of power consumption, the overall rack power consumption of Panjiu AL144 is as high as 650 kW, with a maximum support for single GPU power consumption of 3,500 W.
To support such massive power consumption, Panjiu AL144 adopts an 800V vertical power supply architecture, coupled with an 800V PDB and vertical jet microchannel cold plate inside the server. This is a fully liquid-cooled native design.
Almost forgot, Alibaba Cloud also has a Panjiu CXL memory expansion super node.
In traditional servers, the CPU is strongly bound to DDR memory, and memory capacity is limited by slots, which also forms resource silos.
The Panjiu CXL memory expansion super node pools memory resources based on the CXL 3.1 protocol, decoupling memory resources from the CPU. It can build a shared memory pool to solve the memory wall problem in large model KV Cache and database scenarios.
█ Clusters
Above servers/super nodes is the cluster level. Alibaba Cloud's complete large-scale computing hardware cluster system is named Lingjun.
The basic computing hardware unit of Lingjun is the Panjiu AL64/AL128/AL144 super node, which needs no further explanation.
The key to the cluster lies in the Scale-out network, which is the high-speed interconnect network between nodes. The Lingjun cluster adopts a self-developed HPN (High Performance Network) high-performance network architecture, supporting the RDMA protocol. It can achieve linear acceleration ratio at the 10,000-GPU scale, ensuring efficient communication for large-scale distributed training.
HPN has two mainstream versions: 7.0 and 8.0.
HPN 7.0 was deployed on a large scale in 2023, featuring a 51.2T Ethernet switch chip. The NICs are mainly 400G RDMA, capable of building a single-layer 1,000-GPU and two-layer 10,000-GPU cluster. The computing network and storage network are separated into dual planes. The self-developed Solar-RDMA + HPCC flow control + ACCL communication library reduces AI training communication latency and jitter, improving the overall training efficiency of the cluster.
HPN 8.0 was released at the 2025 Apsara Conference, paired with the Zhenwu M890 super node. It supports 64-GPU super node Scale-up, supporting trillion-parameter MoE models like Qwen 3.8, with GPU interconnect bandwidth of 6.4 Tbps and storage network bandwidth of 800 Gbps.
The HPN 8.0 Pro enhanced version can support 130,000 800G ports per cluster, with the cluster upper limit reaching the 100,000-GPU scale.
Alibaba Cloud's self-developed switch seems to have been named Baihu, which remains to be confirmed. It doesn't seem to be written this way in official documents and news; it might be an internal nickname.
For the core ASIC switch chip, Alibaba Cloud showcased a 102.4T product (UPO, Ultra Performance NPO).
The self-developed 51.2T router comes in both 128×400G and 64×800G specifications, which can build connections between data centers (Scale-Across), providing flexible networking options for building larger-scale and higher-bandwidth computing clusters.
By the way, let's mention Alibaba Cloud's optical communication technology level.
Alibaba Cloud is very strong in computing chips, which is still understandable. However, visiting this Apsara Conference, I found that Alibaba Cloud is also astonishing in the field of optical communication.
First, Alibaba Cloud has self-developed 800G/1.6T optical modules. Then, there is also the self-developed C+L single-wavelength 1.6T transmission equipment.
In recent years, the very popular CPO (Co-Packaged Optics) and NPO (Near-Package Optics) technologies in optical communication have been deeply laid out by Alibaba Cloud. For silicon photonics chips, Alibaba Cloud has self-developed ones. Liquid cooling solutions are also under research.
How to put it, in recent years, "Yi Zhongtian" (a popular nickname for leading optical communication companies) has been thriving in the optical communication circle, with stock prices surging (though they dropped recently). Looking at this, Alibaba Cloud's self-developed capabilities are not much inferior. If this capability were spun off to establish an independent company, it would probably also occupy a place in the optical communication market.
You can say that Huawei, with its background in communications, naturally has deep accumulation in optical communication technology. But Alibaba Cloud did not originate from communications; how could it self-develop so many things? Anyway, I am quite impressed.
In terms of cluster storage, Alibaba Cloud has the self-developed Pangu distributed storage system, as well as CPFS (Cloud Parallel File Storage).
This CPFS is the cluster-level shared storage supporting the Lingjun computing cluster, capable of meeting the high-throughput and low-latency data access requirements in large-scale AI training scenarios.
█ Conclusion
Alright, the above is the inventory and combing of Alibaba Cloud's AI computing hardware product system. Some edge-side products like Wuying, as well as some software ecosystems and platforms like Zhenwu PPU SDK and PAI, I will not elaborate on.
In summary, as a cloud service provider, Alibaba Cloud's self-developed computing system has achieved a full-stack closed loop from underlying chips to upper-layer clusters, and its technical level has entered the global first tier. This is indeed very surprising. It feels they have been too low-key usually; only when they suddenly reveal their strengths do people realize the depth of their technical reserves.
As mentioned before, Alibaba Cloud's "self-production and self-consumption" model has both advantages and disadvantages. Frankly speaking, Xiao Zao Jun actually prefers specialized division of labor. Spinning off product capabilities independently, incubating them into independent commercial entities, facing more clients, and thereby forming a broader domestic ecosystem might be better.
Of course, this is just my personal opinion. Computing power, storage power, and cluster communication constitute a huge market, and domestic vendors have a lot of room for development. How Alibaba Cloud's unique model will develop remains to be tested by time.