Hardware basics
A real AI server doesn't run on just one kind of chip. A couple of CPUs orchestrate, many GPUs — or TPUs, on Google's own stack — do the heavy math, and a DPU quietly keeps data flowing so nothing stalls. These aren't competing chips; they're teammates with very different jobs.
CPU
The general-purpose brain of almost every computer. It runs the operating system, handles input and output, and executes instructions one after another extremely fast — built to do any task well, including ones where each step depends on the result of the last.
GPU
Originally built to render graphics, which means doing the same simple calculation across millions of pixels at once. That same trick turned out to be exactly what AI needs — repeating the same math, like matrix multiplication, across huge amounts of data in parallel — which is why GPUs ended up running the field instead of graphics cards.
CPU vs. GPU
The real difference: a CPU is built around a handful of large, complex cores meant to handle any task well. A GPU is built around thousands of small, simple cores meant to repeat the same math across huge amounts of data at once. That's also why GPUs get ultra-fast HBM memory stacked directly on the chip: feeding thousands of cores at once takes far more bandwidth than a CPU's smaller core count ever needs.
TPU
A TPU (Tensor Processing Unit) takes specialization even further than a GPU. Where a GPU is a general-purpose parallel processor that happens to be great at AI math, a TPU is an ASIC — a chip hard-wired for one job, the matrix multiplication at the core of neural networks. That trade cuts flexibility but improves efficiency per watt. TPUs are Google's own chip, used internally and rented out through Google Cloud; Ironwood (TPU v7) was built primarily for inference at scale, and Google has since split its 8th-generation TPUs into separate training (TPU 8T) and inference (TPU 8I) chips. Amazon (Trainium) and Microsoft (Maia) have since built their own equivalents, for the same reason: less dependence on Nvidia and lower cost per inference.
DPU
A DPU (Data Processing Unit) does something different from all three chips above — it doesn't do AI math at all. It's a specialized chip that offloads networking, storage, and security tasks away from the CPU. That matters in an AI cluster because training and inference move enormous amounts of data between GPUs constantly; if the CPU has to manage that network traffic itself, it becomes a bottleneck, leaving the far more expensive GPUs sitting idle waiting for data. Nvidia's BlueField is the best-known example, with AMD's Pensando as another option.