Hardware supply chain, six stages
Every AI model runs on hardware that starts as raw sand and ends as a building that draws as much power as a small city. Here's what happens at each stage between those two points — what it is, who builds it, and what it costs.
This is the smallest, simplest thing in the whole chain — the raw material every GPU is ultimately made of. Every stage after this one is just this disc getting cut up, built on, and multiplied.
A cylinder of ultra-pure silicon crystal, sliced into thin polished discs — 300mm across at leading fabs. This disc is the substrate every transistor on Earth gets built on.
Before it can be used as a die, the blank disc goes through hundreds of repeated steps — lithography, etching, and doping — that build transistors and wiring into it, layer by layer. Each extra step adds cost, which is why a blank wafer runs a few hundred dollars but a fully processed one can run into the tens of thousands.
+ lithography, etch & doping — hundreds of dies cut from one wafer, then tested and binned
A wafer on its own isn't useful — it gets cut into a grid of identical squares, and each square is a die. One die is the actual compute part of what eventually becomes a GPU.
A standard 300mm wafer typically yields somewhere between 100 and 300 of these dies, depending on how large each die is and how many are lost to defects. Each square on the wafer gets etched with billions of tiny transistors, and a modern GPU die packs tens of billions of them into an area smaller than a postage stamp.
Once the wafer is fully patterned, it gets diced into individual dies, then each one is electrically tested and sorted — or "binned" — by which came out fully functional. Not every die on a wafer works, and the yield loss from the ones that fail is part of what the working dies end up costing.
+ HBM memory stacks + advanced packaging — a bare die becomes a sellable GPU
A bare die can't do anything on its own — it needs memory sitting right next to it to actually process data. Add that memory and package it up, and a die becomes what most people actually mean by "GPU": a chip you can buy, like an H100.
Manufacturers stack High Bandwidth Memory chips next to the die on a silicon interposer and package the whole assembly — this is what turns a die into a named product like the H100, H200, or B200.
Getting from a bare die to a sellable chip means bonding the die and its HBM stacks onto a silicon interposer — advanced packaging, done mostly by TSMC — then testing and binning the finished part. Chips that clock higher or have more working memory get sorted into the pricier tiers.
+ ×8 GPUs, host CPUs, and NVLink — assembled onto one board
One GPU isn't powerful enough on its own for today's AI models. Eight get wired together onto a single board so they can share memory at very high speed and act more like one much bigger chip.
Eight GPUs, host CPUs (Nvidia's own Grace chip, in Nvidia's own systems), fast networking cards, and NVLink connections between every GPU are assembled onto one chassis.
Turning eight individual GPUs into one server means mounting them with host CPUs and networking cards onto a single board, wiring NVLink between every GPU pair, then running the whole assembled system through burn-in testing before it ships.
+ ×18 trays + a switch fabric — racked so all GPUs act as one
One server still isn't enough. Dozens of servers get stacked and wired into a rack so every GPU inside it can talk to every other GPU as if the whole rack were a single, much bigger computer.
Multiple compute trays plus dedicated NVLink switch trays combine into one rack. Nvidia's current flagship, the GB200 NVL72, packs 72 GPUs and 36 Grace CPUs into one liquid-cooled rack.
Building a rack means slotting multiple compute trays and dedicated NVLink switch trays into one frame, wiring them together, plumbing in liquid cooling, then testing the whole assembly as a single unit before it ever reaches a data center.
+ ×thousands of racks + a power substation and cooling plant
A rack full of GPUs draws enormous power and heat. This is the final stage: thousands of racks housed in a purpose-built building with its own power and cooling — a facility that can draw as much electricity as a small city.
Thousands of racks installed in a purpose-built facility with dedicated power substations, liquid-cooling plants, and networking that ties every rack together — turning individual racks into one enormous compute cluster.
Standing one up means building the facility itself, connecting it to a dedicated power substation, installing the cooling plant, then racking and cabling thousands of individual racks together. Commissioning and testing the power and cooling systems is often what takes longest before a site goes live.
The same handful of company types show up at every stage.
Fabless designers
Design the chip, never touch a factory.
Nvidia, AMD, Google TPU, AWS Trainium, Microsoft Maia
Foundries
Manufacture chips designed by others.
TSMC, Samsung Foundry, Intel Foundry, GlobalFoundries
Equipment & materials
Build the machines and wafers foundries need.
ASML (lithography), Shin-Etsu, SUMCO (wafers)
Memory makers
Supply the HBM stacked next to every GPU die.
SK Hynix, Samsung, Micron
OEM integrators
Build the physical servers and racks.
Dell, Supermicro, Foxconn, Quanta, HPE
Hyperscalers & neoclouds
Build or rent out the finished data centers.
Microsoft, Meta, Google, Amazon, CoreWeave, Together AI