EN / 中文

The Breaking Point of AI Data Center Investments: GPU, HBM and Power Cost Crisis for Hyperscalers

by bandaotichanyezongheng·June 4, 2026

This article will reverse-engineer the "breaking point of AI data center investments" based on GPU, HBM, and power costs.

Investments in AI data centers have clearly reached unprecedented levels. Hyperscale data center giants such as Microsoft, Google, Amazon, and Meta are competing to invest hundreds of billions of dollars annually. According to TrendForce, by 2026, the total investment from these top four hyperscalers will reach as high as $755 billion (Figure 1). Calculated at an exchange rate of 160 yen to 1 US dollar, this is equivalent to approximately 120.8 trillion yen, exceeding Japan's national budget for fiscal year 2025 (the total general account budget is approximately 115 trillion yen, source: Ministry of Finance of Japan).

Figure 1: The Frenzy of Capital Investment in Data Centers by the Top 4 Hyperscale Operators

The reason for such massive investment is the skyrocketing prices of AI semiconductors used in AI servers. Taking the GPUs of NVIDIA, a leading AI semiconductor manufacturer, as an example, in its current flagship architecture "Blackwell," the price of a single "B200" GPU ranges between 5 million and 8 million yen, a "DGX B200" server equipped with 8 B200 GPUs costs between 40 million and 70 million yen, and the AI rack based on this server reaches up to hundreds of millions to 1 billion yen (Figure 2). Since building AI data centers requires the massive deployment of these AI racks, the investment for each hyperscale data center operator exceeds $100 billion to $200 billion.

Figure 2: Pricing Structure of NVIDIA GPU AI Servers and Data Centers (Hopper, Blackwell, Rubin)

However, this has gone beyond what the term "growth investment" can explain, resembling more like "military buildup for competition."

Under these circumstances, there is a critical question that is rarely discussed directly: "Can this investment really recoup its costs?" While the AI boom emphasizes strong demand and technological innovation, for a capital-intensive industry, the ultimate question is whether the investment can pay for itself.

This article breaks down the cost structure of AI data centers into three elements: GPUs, High Bandwidth Memory (HBM), and power. Furthermore, utilizing publicly available actual data from Microsoft and Google, this article conducts a quantitative analysis of the revenue structure of current AI investments. Based on this analysis, the article attempts to estimate the "breaking line," which is the critical point where investments cannot be recouped.

Please note that this analysis focuses on the direct revenue generated from hourly billing of GPU infrastructure and does not include indirect revenues brought by AI (such as improved search ad quality or increased SaaS value). Please keep this in mind while reading this article.

To put it bluntly, the seemingly frantic investments by US hyperscale data center operators in AI data centers are likely already doomed to fail. Borrowing a famous quote from Kenshiro in the anime "Fist of the North Star": "You are already dead."

The Reality of Investment Scale Seen from the Cases of Microsoft and Google

Figure 3 quantitatively displays the actual investment scale of Microsoft and Google. Based on this data, it fully illustrates that the investment scale of Microsoft and Google (a subsidiary of Alphabet) in the data center field is remarkably massive.

Figure 3: Actual Investment Scale of Microsoft and Google

The Case of Microsoft

According to Microsoft's fiscal year 2025 annual report, capital expenditures (excluding fixed assets and equipment) are expected to reach $64.5 billion. Additionally, the company stated that investments (primarily for AI infrastructure) are projected to exceed $80 billion.

Compared to Microsoft's cloud business revenue of $168 billion, capital expenditures account for approximately 38% of revenue, or about 48% according to the company's statement. Typically, in stable infrastructure businesses, capital expenditures rarely exceed 30% of revenue, making this ratio extremely rare.

More importantly, depreciation expenses have reached $22 billion. This means the burden of past investments has begun to affect the company's profit and loss, and this burden is likely to continue increasing in the coming years. Furthermore, as shown in Figure 1 above, Microsoft's capital expenditure in 2026 is projected to reach $190 billion, approximately 2.4 times that of the previous year. Therefore, Microsoft's profitability is expected to decline significantly.

The Case of Google

Meanwhile, Google's parent company, Alphabet, is making even larger-scale investments. Its capital expenditure in 2025 reached $91.4 billion, most of which will be used for technical infrastructure such as servers and data centers. In contrast, Google Cloud's annual revenue is approximately $58.8 billion, with an operating profit of about $13.9 billion.

Of course, this $91.4 billion in capital expenditures supports not only the cloud computing business but also company-wide infrastructure, such as search engines and AI research platforms. However, even if half of it is used for cloud computing services, it still amounts to approximately $45.7 billion, which is about 80% of cloud computing sales and roughly 3.3 times the operating profit. Even considering this, it is evident that the current investment scale significantly deviates from traditional return models.

Furthermore, similar to Microsoft, Google's overall capital expenditure in 2026 is projected to reach $180 billion to $190 billion, approximately 2.4 to 2.5 times that of the previous year. Given such a high level of capital expenditure, it is not hard to imagine that recouping the investment in the cloud computing business will become more difficult.

The Cost Structure of AI Data Centers

The reason for this massive investment lies in the unique cost structure of AI data centers. First, we will estimate the cost structure and market scope of AI data centers (Figure 4).

Figure 4: Cost Structure and Market Scope of AI Data Centers

First, let's look at GPUs. Current AI infrastructure relies almost entirely on NVIDIA's GPUs. For example, the unit price of an H100 system is estimated to be between $25,000 and $40,000, depending on the configuration, while a server rack containing 8 H100s will cost approximately $3 million. Furthermore, the price of GB200 series racks is expected to rise to several million dollars (approximately $3.5 million to $5.5 million).

Another crucial factor is that the investment target is not a single GPU, but a "cluster unit." In current AI data centers, it is common to deploy thousands to tens of thousands of GPUs per cluster, with the investment for a single cluster ranging from hundreds of millions to about $700 million.

Next is HBM (High Bandwidth Memory). In H100 and GB200 chips, each GPU is typically equipped with 6 to 8 HBM stacks. The unit price of HBM varies by generation and contract terms, but the unit price of HBM3/3E is said to be between $1,000 and $1,500. Therefore, the HBM cost per GPU is approximately $10,000, which accounts for a significant proportion of the GPU price.

More importantly, there are supply constraints. The HBM market is almost entirely dominated by three companies: SK Hynix, Samsung Electronics, and Micron Technology. In particular, SK Hynix is said to hold over 50% of the advanced HBM market. This concentration of supply creates a structure that suppresses price declines.

Third, there is the power consumption issue. The power consumption of AI data centers is several orders of magnitude higher than that of traditional cloud platforms (Figure 5). For example, the TDP (Thermal Design Power, referring to the estimated maximum heat generation required to cool the chip) of the H100 is about 700W, while the TDP of the GB200 is at the 1kW level. If a cluster containing 10,000 GPUs is configured, the power consumption of the GPUs alone will reach 10MW, and with other power consumptions such as networking and cooling, the total power consumption will reach 20-30MW.

Figure 5: Annual Power Consumption and Total Cost of AI Data Centers

Returning to the explanation of Figure 5, converted to annual power consumption, a 20-megawatt system requires 20 MW × 24 hours × 365 days ≈ 175 million kWh/year. Assuming an electricity price of $0.14/kWh, the annual electricity cost is approximately $25 million. In practice, considering redundant configurations and cooling losses, it is not uncommon for costs to reach around $35 million per year.

Therefore, these three elements—GPUs (CapEx), HBM (supply constraints), and power (OpEx)—all grow exponentially as scale expands. As a result, the cost of AI infrastructure remains persistently high, and it seems difficult to reduce costs through scale expansion as in the past.

Traditional Recovery Models Are Not Feasible

Traditional cloud infrastructure benefits from economies of scale, driven by the continuous decline in server unit costs and improved utilization rates. Advances in Moore's Law and virtualization technology have enabled a single server to "handle more services at a lower cost" over time, supporting the recovery model. However, the situation in AI data centers is entirely different. Figure 6 shows the prerequisites for its cost structure, and Figure 7 shows the recovery line of AI data centers calculated based on these conditions.

Figure 6: Assumptions for AI Data Center Recovery Model Calculation

Figure 7: AI Data Center Payback Period Calculation

Assuming an initial investment of $700 million to build a cluster with 10,000 GPUs (including GPUs, servers, networking, and cooling systems), and amortizing it over 5 years for accounting purposes, the annual amortization expense is $140 million. Adding $35 million in power costs and $35 million in operating costs (maintenance, personnel costs, data center rent, etc.), the total annual cost is approximately $210 million.

From this, the required billing cost per GPU for recovery can be expressed by the following formula.

Required Billing Cost = Total Annual Cost ÷ (Number of GPUs × 8,760 Hours × Uptime Rate)

Assuming an uptime rate of 70%, $210 million ÷ (10,000 × 8,760 hours × 0.7) ≈ approximately $3.43/GPU hour

In other words, unless each GPU generates at least $3.43 in revenue per hour under near-constant operating conditions, the investment cannot be recouped. This is the "floor," not the "average"; if the utilization rate drops, the required unit cost will be even higher.

However, in the actual market, the prices for generative AI inference are dropping rapidly. For instance, it is reported that the Application Programming Interface (API) prices for Large Language Models (LLMs) will drop to less than one-tenth of their original prices between 2023 and 2025. Furthermore, the proliferation of open-source models has further intensified price competition.

The key point is that despite the significant drop in API prices, the costs of GPUs, HBM, and power are actually rising. At this point, the traditional recovery model is no longer feasible. AI infrastructure is shifting from a model where "the larger the scale, the more obvious the advantages" to a model where "the larger the scale, the higher the fixed cost risk." Then, at what scale does recovery become impossible? Let us analyze the recovery conditions based on the real data from Microsoft and Google.

The Reality of the Recovery Line

As mentioned earlier, Microsoft continues to invest $60 billion to $80 billion annually, and by 2025, its depreciation expenses have exceeded $20 billion. If Microsoft attempts to cover this $22 billion in depreciation expenses with the operating profit of Microsoft Cloud, it will significantly reduce the operating profit margin of its cloud business. On the other hand, Google Cloud's operating profit is $13.9 billion, while its capital expenditure for the cloud business alone is as high as approximately $45.7 billion, meaning that even on a single-year basis, its investment exceeds three times its operating profit.

This indicates the existence of structural problems. AI infrastructure must maintain an extremely high return on investment to be profitable. However, the reality is that the prices of AI services are falling, the costs of GPUs and HBM remain high, and power costs are rising.

In an environment where these three factors act simultaneously, the conditions for investment recovery will deteriorate rapidly. It can be said that current AI investments have entered a structural dilemma: it is difficult to recoup investments unless extremely high utilization rates and extremely high unit prices are achieved simultaneously.

Why Investments Must Continue

So, will this frantic investment in capital equipment slow down? The answer is no.

Microsoft's remaining performance obligations are approximately $368 billion, indicating that market demand still exceeds supply. Google has also explicitly stated its plan to further expand capital expenditures to meet the demands of AI and cloud computing. The key point is that neither company is investing because they expect to recoup the investment. Instead, they are forced to continue investing because stopping investment means falling behind in the competition.

Current AI investments have shifted from pursuing profit maximization to striving to avoid failure. We should view AI investments as having entered a "war of attrition" stage rather than a "growth" stage.

As long as this structure persists, the AI boom will continue to expand, but internally it will accumulate a "distortion" in the form of irrecoverable risk. This distortion will suddenly manifest at a certain node. This is the "breaking line" that will be elaborated in the next chapter.

Exploring the Breaking Line

As mentioned above, judging the sustainability of AI investments requires considering not only the number of GPUs but also HBM, power, and the entire power infrastructure. This article will take a cluster with 10,000 GPUs as an example to quantitatively demonstrate at what scale investment recovery will become impossible—the so-called "breaking line."

Working backward from the number of GPUs, HBM and power consumption increase as follows:

First, we assume a cluster composed of 10,000 GPUs. Figure 8 shows the required annual power consumption for each cluster and the required number of equivalent nuclear power plants.

Figure 8: Physical Scale of Power Consumption Required for the Breaking Line

Assuming each GPU is equipped with 8 HBM stacks, the total required HBM will reach 80,000 stacks. With 24GB per stack, the total is approximately 1.92PB. Furthermore, in terms of power consumption, assuming each GPU consumes 1kW, and the entire facility (including cooling, substations, and network loads) consumes about twice that amount, the facility load for a cluster with 10,000 GPUs is approximately 20MW.

The annual power consumption is approximately 175.2 gigawatt-hours (GWh). Divided by the annual power generation of a 1-gigawatt-class nuclear power plant operating at 90% load, this is equivalent to the power generation of about 0.022 reactors. Conversely, this means a single nuclear power plant can only meet the electricity needs of about 45 such sites. If AI clusters expand on a large scale, the demand cannot be met without building new nuclear power plants.

Definition of the Breaking Line

As mentioned above, assuming a cluster with 10,000 GPUs, an initial investment of $700 million amortized over 5 years, annual operating costs of $35 million, and annual power costs of approximately $35 million, the total annual cost is about $210 million. Under these circumstances, the break-even condition can be expressed by the following formula described in Chapter 3.

Required Billing Cost = Total Annual Cost ÷ (Number of GPUs × 24 Hours × 365 Days × Uptime Rate)

Assuming a utilization rate of 70%, the billing cost per GPU hour is approximately $3.43. This article refers to this as the "critical point." In other words, once the price of AI services falls below this level, or the utilization rate drops below this assumed value, the investment cannot recoup its costs.

It should be noted that the 5-year amortization period adopted for accounting purposes is a relatively optimistic assumption compared to the technology cycle of NVIDIA GPUs, which typically updates every two years or so. In the breaking scenario ③ described later, we will analyze the impact of this shortened amortization period on the revenue structure.

The Break Happens Suddenly

In typical infrastructure industries, profit margins gradually decline. However, in AI data centers with extremely high fixed costs, once the profit margin falls below a certain level, profitability deteriorates rapidly due to the following three reasons.

First, the initial investments in GPUs and HBM are massive and fixed.

Second, power and cooling loads are high and not easily reduced.

Third, on the other hand, due to competition, the required billing unit price (market price) will decline.

Therefore, the deterioration process of AI investments is not linear but non-linear. In other words, it is not "the situation gradually worsens and then becomes more difficult," but rather "once a certain critical point is crossed, losses suddenly become massive." This is the essence of the breaking line.

Now, let us quantitatively calculate three scenarios where AI data centers break down. The common conditions for each scenario are shown in Figure 9.

Figure 9: Common Conditions for Calculating the Breaking Line of AI Data Centers

Three Breaking Scenarios

Figure 10 shows the simulation results of the three breaking scenarios.

Figure 10: Simulation of Three Breaking Scenarios for AI Data Centers

First, Software Break.

The most likely scenario is fierce price competition among AI companies. If the billing price drops to $2.90 per GPU hour and the utilization rate drops to 65%, the required billing price will rise to $3.69, resulting in an annual loss of approximately $44.9 million. However, as shown in Figure 10, although a complete break has not occurred at this stage, profits have completely disappeared, and investment recovery is quietly heading towards failure. Even if surface demand is maintained, internal capital efficiency is plummeting.

Second, Hardware Break.

The next risk is the rise in actual costs such as power, cooling, and installation. With a billing rate of $3 and a utilization rate of 55%, coupled with rising electricity prices and increased facility loads, the required billing rate will jump to $4.7, resulting in an annual loss of about $81.7 million. Figure 10 shows that at this stage, the deficit expands sharply. This is a typical example of how infrastructure costs, rather than demand, destroy profitability.

Third, Financial Break.

The most severe consequence is a financial break. Even if the billing rate is $3.20 per unit and the occupancy rate is 60%, due to the shortened depreciation period (from 5 years to 4 years) and an 8% cost of capital, the actual billing rate needs to reach $5.73 per unit, resulting in an annual loss of approximately $133 million. Therefore, as shown in the bottom row of Figure 10, the losses at this stage have reached an unbearable level ($133 million per year). The essence of this situation is that the capital market determines the investment is "irrecoverable" before the equipment physically fails.

Failure Occurs in a "Non-linear" Manner

Figure 11 shows the relationship between AI data center utilization rate and required billing cost. It should be noted that this relationship is not linear.

Figure 11: The Zone Where AI Data Centers Will Break Down

When the occupancy rate is 70%, the required unit cost is approximately $3.43; but when the occupancy rate drops to 60%, the required unit cost will rise to nearly $4; if the occupancy rate further drops to 50%, the required unit cost will jump to nearly $5.

The "breaking zone" shown in Figure 11 intuitively demonstrates this non-linear relationship. The market price range ($2.5 to $3.0: the hourly rate range for H100/H200 based on platforms such as AWS, Azure, Lambda Labs, etc.) has already fallen deep into this zone, and it is highly likely that current AI service prices are structurally below the break-even point.

Power Consumption Limits: AI Is a National Infrastructure Issue

More importantly, the scaling of AI investments directly depends on power infrastructure. As shown in Figure 12, 10,000 GPUs require approximately 20 megawatts (MW) of power, 100,000 GPUs require 200 megawatts (MW), and 1,000,000 GPUs require 2,000 megawatts (MW) (= 2 gigawatts (GW)). This means that not only do data centers need to be expanded, but the power supply infrastructure itself must also be expanded.

Figure 12: Power Consumption Increases Sharply from 10,000 GPUs to 100,000 GPUs and Then to 1,000,000 GPUs

If we convert this power into nuclear energy:

Cluster of 10,000 GPUs: 0.02 units

Cluster of 100,000 GPUs: 0.2 units

Cluster of 1,000,000 GPUs: 2.2 units

The expansion of AI investments is clearly equivalent to the expansion of power infrastructure. AI data centers are no longer just an issue for the IT industry but have evolved into a "national supply capacity issue" involving power, land, and construction capabilities.

The "Break" Facing AI Investments

Current investments in AI data centers are not only unprofitable but also physically unsustainable. Falling market prices, declining utilization rates, rising power costs, or tightening capital markets—any single one of these factors could immediately drive data centers to the breaking point. Moreover, this break will not happen gradually but will erupt suddenly once a certain critical point is crossed. This is no longer just an issue for the semiconductor industry but a matter concerning national power supply capacity.

On April 3, 2026, Japanese Minister of State for Economic Security Sanae Takaichi met with Brad Smith, President of Microsoft, a major US hyperscale data center operator, and welcomed the company's investment of approximately $10 billion in data centers in Japan. However, as this article shows, such investments are not only unprofitable but also consume massive amounts of power, and their structure will burden national infrastructure. Behind the AI boom, it is necessary to calmly assess the scale of the price Japan will have to pay.

*Disclaimer: This article is created by the original author. The content represents their personal views. We repost it solely for sharing and discussion, which does not mean we endorse or agree with it. If you have any objections, please contact the backend.