Stanford MS&E435: The Scarce Megawatt Behind the AI Factory
After this lecture, “cloud AI” looks remarkably like heavy industry: electricity, substations, cooling, concrete, and the people assembling it all. That is what makes the third Stanford MS&E435 session useful—it takes a conversation about models, tokens, and GPUs down to the physical layer beneath them. The guest is Chase Lochmiller, co-founder and CEO of Crusoe, which builds and sells the entire stack from site development to managed inference. This makes the session both a useful insider map and a vendor pitch: Lochmiller explains the same vertically integrated stack his company sells.
His framing is “from electrons to tokens”: energy → building and cooling → compute → tokens → completed work. His central claim, based on Crusoe's experience, is that the current bottleneck is no longer an individual GPU but an energised data centre—a place where the chip can actually be powered on and put to work.
These are the insights I would take from the lecture:
🔸 The bottleneck keeps moving Yesterday it was GPUs; today it is connected power, gas turbines, transformers, switchgear, and qualified electricians. Vertical integration is not only a way to capture more margin but also insurance against the next shortage.
🔸 The scarce product is an energised megawatt Buying accelerators is not enough: a project needs land, permits, a substation, networking, and cooling. A site that can bring 100–1,000 MW of compute online quickly therefore becomes a strategic asset.
🔸 Moving data can be cheaper than moving electricity Training does not have to sit close to the user, so Crusoe places clusters near surplus energy in West Texas. The logic is weaker for latency-sensitive inference and data with strict residency requirements.
🔸 The AI boom is also an industrial boom According to Lochmiller, thousands of construction workers are on the Abilene campus every day. Electricians, welders, pipefitters, concrete, and kilometres of cable become as important as CUDA and model architecture.
🔸 An old GPU can become a commodity; scale does not A large coherent cluster connected by a shared high-performance network is difficult to reproduce. A managed service can also hide the hardware generation from the customer and extend an accelerator's economic life: the customer needs an API result, not a particular H100.
🔸 In Lochmiller's model, the economics depend on depreciation, utilisation, and the service layer He estimates the building and generation layer at about $20 million per MW and IT at another $40 million, roughly $30 million of which goes to GPUs. Bare compute rental generates around $15 million per MW annually in his calculation; managed inference could reach $30 million in an optimistic case. I would treat those figures only as the CEO's own model. Lochmiller explicitly acknowledges possible double counting, incomplete operating expenses, and the fact that the slides were assembled that same afternoon. The advertised two-to-four-year “payback” is therefore closer to CapEx divided by revenue than a profit or free-cash-flow calculation.
The most debatable part is the final arrow in the “electrons to tokens” chain: from tokens to completed work. Lochmiller calls AI agents digital labour and maps them onto the labour input of a Cobb–Douglas model. It is a production function (or utility function) that expresses output Y as a function of its production factors—the labour input L and physical capital K). It is written as Y(L,K)=A * (L * alpha) * (K * beta), where A is the technology coefficient, α ≥ 0 is the output elasticity of labour, and β ≥ 0 is the output elasticity of capital. It works as a metaphor. In conventional growth accounting, however, GPUs and data centres look more like capital, model improvements like technology and productivity, while tokens by themselves are not yet the same thing as useful output.
One more caveat: closed-loop cooling in Abilene does sharply reduce water consumption, but Lochmiller says a 350 MW gas plant was built for the campus. Low-water does not automatically mean low-carbon.
I would recommend watching 9:26–18:30 and 23:54–40:45: the first section changes the picture of the bottleneck; the second offers a rare public breakdown of megawatt economics, complete with both useful numbers and unusually candid caveats.
#AI #Engineering #Infrastructure #Architecture #Economics #Bigtech
Public sources
- Stanford Online: MS&E435 — Building AI Factories
- Stanford MS&E435: course schedule and materials
- Crusoe: Chase Lochmiller biography
- Crusoe: projected expansion of the Abilene campus to 2.1 GW
- IEA: data-centre electricity growth and physical bottlenecks
- OpenAI: Abilene infrastructure and water cooling
- OECD: AI, capital deepening and multifactor productivity