← Home

Hardware supply chain, six stages

Silicon to Supercluster

Every AI model runs on hardware that starts as raw sand and ends as a building that draws as much power as a small city. Here's what happens at each stage between those two points — what it is, who builds it, and what it costs.

1

Silicon wafer

$3,000 – $30,000 / wafer
An ingot is sliced into wafers; each wafer is patterned into hundreds of identical dies.
Photo of a real 12-inch silicon wafer
A real 12–inch (300mm) wafer. Wikimedia Commons
Why it matters

This is the smallest, simplest thing in the whole chain — the raw material every GPU is ultimately made of. Every stage after this one is just this disc getting cut up, built on, and multiplied.

What it is

A cylinder of ultra-pure silicon crystal, sliced into thin polished discs — 300mm across at leading fabs. This disc is the substrate every transistor on Earth gets built on.

The process

Before it can be used as a die, the blank disc goes through hundreds of repeated steps — lithography, etching, and doping — that build transistors and wiring into it, layer by layer. Each extra step adds cost, which is why a blank wafer runs a few hundred dollars but a fully processed one can run into the tens of thousands.

Who
TSMC (TSM) — make Samsung Foundry (005930.KS) — make ASML (ASML) — supply
Price
Blank — $200–500 28nm node — ~$3,000 2–3nm node — $20,000–$30,000

+ lithography, etch & doping — hundreds of dies cut from one wafer, then tested and binned

2

The chip die

~$100s – low $1,000s / die
compute die + HBM, on an interposer
The bare die is only the beginning — it still needs memory and a package.
Photo of a single bare silicon die
A single bare die. Wikimedia Commons
Why it matters

A wafer on its own isn't useful — it gets cut into a grid of identical squares, and each square is a die. One die is the actual compute part of what eventually becomes a GPU.

What it is

A standard 300mm wafer typically yields somewhere between 100 and 300 of these dies, depending on how large each die is and how many are lost to defects. Each square on the wafer gets etched with billions of tiny transistors, and a modern GPU die packs tens of billions of them into an area smaller than a postage stamp.

The process

Once the wafer is fully patterned, it gets diced into individual dies, then each one is electrically tested and sorted — or "binned" — by which came out fully functional. Not every die on a wafer works, and the yield loss from the ones that fail is part of what the working dies end up costing.

Who
Nvidia (NVDA) — design AMD (AMD) — design TSMC (TSM) — make
Price
Small / older die — $100s Large GPU die — low $1,000s

+ HBM memory stacks + advanced packaging — a bare die becomes a sellable GPU

3

The GPU

$25,000 – $50,000 / chip
GPU
Compute die + HBM + substrate, tested and sold as one packaged part.
Why it matters

A bare die can't do anything on its own — it needs memory sitting right next to it to actually process data. Add that memory and package it up, and a die becomes what most people actually mean by "GPU": a chip you can buy, like an H100.

What it is

Manufacturers stack High Bandwidth Memory chips next to the die on a silicon interposer and package the whole assembly — this is what turns a die into a named product like the H100, H200, or B200.

The process

Getting from a bare die to a sellable chip means bonding the die and its HBM stacks onto a silicon interposer — advanced packaging, done mostly by TSMC — then testing and binning the finished part. Chips that clock higher or have more working memory get sorted into the pricier tiers.

Who
Nvidia (NVDA) — design AMD (AMD) — design TSMC (TSM) — make SK Hynix (000660.KS) — supply
Price
H100 — $25,000–$40,000 B200 — $30,000–$50,000 B300 — $50,000–$60,000

+ ×8 GPUs, host CPUs, and NVLink — assembled onto one board

4

The server

~$300,000 – $350,000 / server
Eight GPUs, two host CPUs, and an NVLink mesh on a single chassis.
Photo of a Supermicro server motherboard
A Supermicro server board — a real OEM in this stage. (Illustrative, not the specific 8-GPU board.) Wikimedia Commons
Why it matters

One GPU isn't powerful enough on its own for today's AI models. Eight get wired together onto a single board so they can share memory at very high speed and act more like one much bigger chip.

What it is

Eight GPUs, host CPUs (Nvidia's own Grace chip, in Nvidia's own systems), fast networking cards, and NVLink connections between every GPU are assembled onto one chassis.

The process

Turning eight individual GPUs into one server means mounting them with host CPUs and networking cards onto a single board, wiring NVLink between every GPU pair, then running the whole assembled system through burn-in testing before it ships.

Who
Nvidia (NVDA) — design Dell (DELL) — make Supermicro (SMCI) — make
Price
Base config — ~$300,000 Fully loaded — up to $350,000

+ ×18 trays + a switch fabric — racked so all GPUs act as one

5

The rack

~$3M – $8.8M / rack
Nvidia's GB200 NVL72: 72 GPUs and 36 CPUs, wired as one liquid-cooled unit.
Photo of a server rack
A server rack. (Illustrative, not the specific GB200 NVL72.) Wikimedia Commons
Why it matters

One server still isn't enough. Dozens of servers get stacked and wired into a rack so every GPU inside it can talk to every other GPU as if the whole rack were a single, much bigger computer.

What it is

Multiple compute trays plus dedicated NVLink switch trays combine into one rack. Nvidia's current flagship, the GB200 NVL72, packs 72 GPUs and 36 Grace CPUs into one liquid-cooled rack.

The process

Building a rack means slotting multiple compute trays and dedicated NVLink switch trays into one frame, wiring them together, plumbing in liquid cooling, then testing the whole assembly as a single unit before it ever reaches a data center.

Who
Nvidia (NVDA) — design Dell (DELL) — make Foxconn (2317.TW) — make
Price
GB200 NVL72 (now) — ~$3M Vera Rubin NVL72 (next-gen) — up to $8.8M

+ ×thousands of racks + a power substation and cooling plant

6

The data center

$30M – $45M / megawatt
power + cooling
Thousands of racks, fed by dedicated substations and liquid-cooling plants.
Photo of the inside of a real data center, showing rows of server racks
Inside a real data center (CERN). Wikimedia Commons
Why it matters

A rack full of GPUs draws enormous power and heat. This is the final stage: thousands of racks housed in a purpose-built building with its own power and cooling — a facility that can draw as much electricity as a small city.

What it is

Thousands of racks installed in a purpose-built facility with dedicated power substations, liquid-cooling plants, and networking that ties every rack together — turning individual racks into one enormous compute cluster.

The process

Standing one up means building the facility itself, connecting it to a dedicated power substation, installing the cooling plant, then racking and cabling thousands of individual racks together. Commissioning and testing the power and cooling systems is often what takes longest before a site goes live.

Who
Microsoft (MSFT) — operate Meta (META) — operate CoreWeave (CRWV) — operate
Price
Per megawatt — $30M–$45M Meta Hyperion — $50B–$60B MSFT Fairwater — ~$7.3B

Who's who in the chain

The same handful of company types show up at every stage.

Fabless designers

Design the chip, never touch a factory.

Nvidia, AMD, Google TPU, AWS Trainium, Microsoft Maia

Foundries

Manufacture chips designed by others.

TSMC, Samsung Foundry, Intel Foundry, GlobalFoundries

Equipment & materials

Build the machines and wafers foundries need.

ASML (lithography), Shin-Etsu, SUMCO (wafers)

Memory makers

Supply the HBM stacked next to every GPU die.

SK Hynix, Samsung, Micron

OEM integrators

Build the physical servers and racks.

Dell, Supermicro, Foxconn, Quanta, HPE

Hyperscalers & neoclouds

Build or rent out the finished data centers.

Microsoft, Meta, Google, Amazon, CoreWeave, Together AI