EN / 中文

From Air Cooling to Immersion Cooling: Three Thermal Routes Tailored for Different Rack Power Densities

by quanqiubandaotiguancha·October 8, 2026

Liquid Cooling Technology

"The Sugon 8000 100,000-GPU AI supercomputing cluster, running at full capacity for a year and a half, generates enough total heat to boil the entire water of West Lake." This widely circulated metaphor makes the thermal load of top-tier AI computing power tangible.

Chip computing power, model parameters, and cluster scales are continuously breaking records every year, yet the physical reality of heat remains unchanged: electrical energy entering a chip is almost entirely converted into heat. The upper limit of computing power ultimately hits a bottleneck in heat dissipation.

Instead of traditional air cooling, Sugon 8000 adopted two-phase immersion liquid cooling combined with diamond-copper composite thermal conductive materials to manage the heat generated by 100,000 chips. This solution is no longer an isolated case in domestic intelligent computing centers.

From the data centers of leading cloud providers to commercial AI servers, heat dissipation has evolved from a mere data center supporting issue into a critical prerequisite for computing power deployment, cost reduction, and stable operation.

01. Computing Power is Surging, but Heat Dissipation Hits the Wall First

In recent years, the power curve of AI clusters has risen rapidly. In the early days, the power consumption per rack in data centers was generally 5-10kW, which could be easily managed by standard air cooling. Currently, the power consumption per rack in AI training clusters generally exceeds 80kW and 100kW, and the rack power consumption of 100,000-GPU clusters is approaching the megawatt level. The power consumption of a single AI chip has also increased from the hundreds of watts to the kilowatts level, with the local heat flux density multiplying.

Power consumption is growing fast, while heat dissipation is lagging behind, making the pain points evident. Under traditional air cooling, high-load servers will trigger chip frequency reduction and computing power shrinkage, increasing the risk of downtime, while the energy consumption for heat dissipation cannot be brought down.

Data disclosed by multiple intelligent computing centers shows that the heat dissipation energy consumption in legacy air-cooled data centers can account for 30%-50% of total electricity usage, making electricity bills the heaviest operational cost. In the early years when the Guizhou National Computing Hub Base used air cooling, the effective computing power of servers was only utilized to 70%, bottlenecked by protective frequency reduction caused by high temperatures.

The large-scale and high-density upgrade of computing power has become the industry norm. Recently, China Mobile Beijing Company, in conjunction with China Mobile Cloud Company, announced that the Huawei Ascend 950DT 1,000-GPU supernode has officially landed at the Mobile Cloud Beijing Node of the Information Harbor Data Center in Changping District, once again confirming the trend of computing clusters iterating towards ultra-high density and ultra-large scale. Based on Huawei's supernode architecture, this supernode is equipped with 1,024 cards of computing power per node, with a peak computing power reaching 1 EFLOPS (FP8). It supports 256TB of unified memory and microsecond-level ultra-low latency, comprehensively supporting high-level computing power needs such as large-scale training and inference of large models, agent development, and the implementation of complex AI scenarios.

The continuous deployment of such high-density computing clusters at the 1,000-GPU and 100,000-GPU levels allows for the continuous expansion of computing power scale. However, it also means that the overall thermal load of the cluster and the local heat flux density of chips are surging simultaneously. The shortcoming of heat dissipation will be further amplified, becoming the core bottleneck restricting the stable operation and performance release of ultra-high-density computing clusters.

02. Three Technical Routes Tailored to Different Power Densities

Currently, the industry is pursuing three parallel routes: air cooling, cold plate liquid cooling, and immersion liquid cooling, each tailored to specific computing power densities.

Air Cooling

Air cooling relies on forced convection from server fans and heat exchange via precision air conditioning in the machine room. With a mature supply chain, simple operation and maintenance, and low retrofit costs, it remains the mainstream for legacy IDCs (Internet Data Centers) and inference clusters of mid-sized cloud providers. However, its shortcomings are also obvious: the heat exchange efficiency of air is low, capping the power per rack at around 30kW. Under high loads, it fails to control noise and PUE (Power Usage Effectiveness), nor can it resolve local hotspots on chips.

New projects in high-density AI scenarios basically no longer choose it, reserving it only as a supplement for low-power inference scenarios. The industry generally believes that under the same AI chip configuration, long-term full load in air-cooled clusters will cause chip temperatures to be 15℃-25℃ higher, accelerating hardware aging and making it difficult to release computing power at full load.

Cold Plate Liquid Cooling

Cold plate liquid cooling attaches microchannel metal cold plates to the surfaces of core heat-generating chips such as GPUs and CPUs. The coolant circulates within the plates to carry away core heat, while components like memory and power supplies still rely on air cooling. It is a compromise solution transitioning from traditional machine rooms to full liquid cooling.

Compared to immersion cooling, it requires minimal server modifications and can leverage traditional operation and maintenance. According to manufacturers' estimates, the initial construction cost can be 30%-40% lower. However, it only covers core chips, resulting in incomplete whole-machine heat dissipation, an upper limit on power consumption per rack, and leakage risks at pipeline joints.

Inspur's public actual tests show that its commercial cold plate liquid cooling solution can control the PUE to within 1.15, covering most general computing and intelligent computing upgrade scenarios. Mid-sized AI computing clusters of Tencent, Baidu, and Huawei Cloud mostly adopt this route.

Compared to air cooling, cold plate liquid cooling can reduce machine room energy consumption by over 25% and lower chip operating temperatures by 10-18℃.

Immersion Cooling

Immersion cooling is the route with the highest heat dissipation efficiency, divided into single-phase and phase-change types.

Single-phase immersion submerges the entire server in dielectric coolant, relying on the sensible heat of the liquid for heat exchange. Phase-change immersion uses low-boiling-point refrigerants, relying on liquid boiling and vaporization to carry away large amounts of heat; Sugon 8000 adopts the phase-change route. It can achieve dead-corner-free heat dissipation for the entire machine and eliminate local hotspots, with the PUE dropping as low as 1.05. The trade-off is that equipment must be redesigned, the coolant is expensive, and the operation and maintenance system is entirely new. Reportedly, high-end fluorinated coolant is quoted at over CNY 500 per liter, and the usage per rack can reach 400-600 liters, making the initial investment considerable.

Immersion cooling is the choice for 100,000-GPU and million-GPU clusters as well as ultra-high-density intelligent computing centers. Alibaba Cloud's Renhe Data Center is one of the largest fully immersion liquid-cooled data centers in the world.

Located in Yuhang, Hangzhou, this project is the first 5A-level green liquid-cooled data center in China. It adopts single-phase immersion liquid cooling technology, submerging the entire server in dielectric coolant, eliminating the need for fans and machine room air conditioning. According to disclosures by Alibaba Cloud, the PUE dropped to a minimum of 1.09 after deployment, saving 70 million kWh of electricity annually, completing the commercial verification of large-scale fully immersion liquid cooling in cloud computing scenarios.

03. The Last Millimeter of Heat Dissipation: Thermal Conductive Materials

Liquid cooling solves "heat exhaust," but the thermal conductive material layer between the chip and the heat sink determines whether heat can be quickly conducted away. If the interfacial thermal resistance is too high, even the best liquid cooling will cause heat accumulation and chip frequency reduction. The main track here is also shifting materials.

Traditional substrates are pure copper and aluminum. Pure copper has a thermal conductivity of 401W/(m·K), with mature processes and low prices, suitable for standard air cooling and low-power cold plates. Once the power consumption of a single chip reaches the kilowatt level and the heat flux density rises, the thermal conductivity rate of copper becomes insufficient, interfacial thermal resistance increases, and heat cannot be conducted away.

Diamond is the material with the best thermal conductivity in nature. Diamond-copper composite materials combine the ultra-high thermal conductivity of diamond with the process compatibility of copper. According to relevant disclosures by Sugon, its thermal conductivity is more than twice that of pure copper. Sugon 8000 directly uses bare chip welding with diamond-copper heat sinks, removing the traditional multi-layer structure of chip metal lids plus thermal pads, and welding the heat sink directly onto the bare chip surface. This compresses the thermal conduction path, reduces interfacial thermal resistance, and solves the problem of local hotspot accumulation on chips in 100,000-GPU clusters, serving as a key support for the long-term full-speed operation of the cluster. Previously, the large-scale implementation of this bare chip direct-welding diamond-copper heat sink process was highly difficult in China.

The thermal interface materials (TIM) between the chip and the heat sink base are also being upgraded. Traditional thermal grease is low-cost, but long-term high-temperature operation causes it to dry out, age, and increase thermal resistance, failing to withstand the 7×24-hour operation of AI servers. According to public reports, high-end AI server models from Huawei and Inspur have batch-replaced thermal grease with liquid metal. New materials such as graphite thermal films and silicon carbide heat sinks are being piloted in edge computing and small high-density servers.

"Boiling the water of West Lake" merely translates the energy consumption of computing power into a perceptible image. AI algorithms can iterate, and computing power scale can double, but the generation and dissipation of heat must follow the laws of physics. From Sugon 8000's phase-change immersion plus diamond heat dissipation solution to the deployment of liquid cooling by various cloud providers and computing hubs, the industry has already invested substantial real money in this path. Whether computing power can continue to scale up and whether the heat dissipation track can keep pace synchronously are among the key variables.

#AI Servers #Liquid Cooling Technology #AI Computing Power