AI Cloud: The Heart of the AI Industry

AI · Industry Structure · Cloud

AI Industry Structure, Layer 3: Cloud

Conceptual image of AI cloud infrastructure
This is the third article in the “5 Layers of the AI Industry” series. If you’re curious about the full structure, we recommend reading The 5 Layers of the AI Industry: Overview first, or if you want Layer 4 (AI Models), check out the AI Models article.

AI model companies like OpenAI and Anthropic often don’t actually build their own servers. Instead, they rent cloud infrastructure. This cloud infrastructure, Layer 3 of the AI industry structure, was a business that existed before the AI boom — but with the arrival of the AI era, it’s become a market where an entirely different scale of money changes hands. In this article, we’ll cover exactly what cloud companies sell, how AWS, Azure, and Google Cloud differ, and why these companies have recently started building their own AI chips.

Advertisement
SECTION 01

Revisiting Cloud

What cloud is
This is the business of building massive data centers packed with tens of thousands of servers, then renting out that computing power as needed. AWS, Azure, and Google Cloud are the leading examples.

Training a single AI model requires tens of thousands of GPUs running non-stop for months. Buying and building all of that directly would take too large an upfront investment, and once training wraps up, all those servers might sit idle. So most AI companies rent cloud like a “subscription service for computing power” — paying only for what they use, and giving it back once they no longer need it.

Cloud itself has actually been a business since before the AI boom. It used to make money from things like website hosting and corporate email servers, but the arrival of the AI era gave it an entirely different scale of revenue.

SECTION 02

How the Big Three Cloud Providers Differ

AWS Azure Google Cloud
Parent Company Amazon Microsoft Google
Market Position #1 by market share #2, strong with enterprise customers #3, strong AI technology
AI-Related Strength Own chip (Trainium) Exclusive OpenAI partnership Pioneer of its own chip (TPU)

All three companies sell the same thing, but each has brought a different weapon to the AI era. Microsoft in particular has invested heavily in OpenAI, effectively turning Azure into ChatGPT’s dedicated infrastructure, while Google is seen as relatively less dependent on NVIDIA thanks to having built its own chip (TPU) for a long time.

SECTION 03

The Rise of Neoclouds

Beyond the big three, a set of companies worth watching has emerged recently — CoreWeave, Lambda Labs, Crusoe, Nebius, and others. The industry calls these companies “neoclouds.” Unlike AWS or Azure, which sell every kind of service from email to databases, these are specialist clouds focused solely on renting out GPUs.

Their weapons are price and speed. With no general-purpose services to worry about, they focus entirely on GPUs, renting out the same performance for far less than the big three. In fact, even OpenAI has become too big for Azure’s capacity alone and rents additional GPUs from CoreWeave — these companies have become impossible to ignore. That said, most of them are still unprofitable and keep expanding their data centers anyway, so it’s worth keeping in mind that this is closer to a bet that “growth will keep going” than a proven track record.

SECTION 04

Why Cloud Companies Are Even Building Their Own Chips

Lately, Amazon, Google, and Microsoft are all building their own AI chips — Amazon’s Trainium, Google’s TPU, and Microsoft’s Maia are the leading examples. This might seem strange — why bother building your own chip when you could just buy NVIDIA’s?

There are two reasons. First, NVIDIA chips are both scarce and expensive, so reducing that dependence can lower costs. Second, an in-house chip can be designed and optimized specifically for a company’s own cloud, leaving room to push efficiency further. That said, these chips still can’t fully replace NVIDIA when it comes to top-tier AI training performance, so most companies take a “mixed strategy,” using NVIDIA chips alongside their own.

SECTION 05

What Investors Should Watch For

The big three cloud providers are currently pouring astronomical capital expenditures (CapEx) into data centers. This has also become one of the hottest debates among investors.

AI cloud revenue growth — How fast AI-related revenue is growing as a share of total cloud revenue
Payback speed on capital spending — How long it takes for the data centers being built right now to actually turn into revenue
Customer concentration — Whether one major customer (e.g., OpenAI) accounts for too large a share of revenue

The scale of cloud companies’ capital spending is the number the market watches most closely every earnings season. It’s a tricky balance to strike: build too much, and worries surface about “overinvestment”; build too little, and worries surface about “falling behind in the AI race.”

Advertisement
CONCLUSION

Closing Thoughts

Cloud looks like it’s just about renting out servers, but for those servers to actually run non-stop, cooling that heat away matters just as much as the servers themselves.

In the next article, we’ll take a close look at the second topic in Layer 3: cooling. We’ll cover why GPU heat has become such a big problem, and why liquid cooling is on the rise.

Next Up — AI Industry Structure, Layer 3: Cooling
We cover why GPU heat suddenly became a problem, plus liquid cooling and the water bottleneck.
Read more →
This document is an overview written from publicly available information about the AI industry’s structure and its leading companies, aimed at helping general readers understand it. Part 4 of the “5 Layers of the AI Industry” series. Last updated: July 2026

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top