AI infrastructure
The hardware that AI runs on: GPUs, the memory and links around them, and the servers, networks and facilities that hold them.
CPU vs GPU
| CPU | GPU | |
|---|---|---|
| Cores | Few, strong cores (tens) | Thousands of simpler cores |
| Built for | Low latency on varied, branching work | High throughput on the same operation applied to lots of data |
| Best at | Running the operating system, control logic, serial code | Matrix math, which is most of deep learning |
Deep learning is mostly large matrix multiplications that split into many independent pieces, which is why GPUs fit it so well. See the labs for how a CPU core handles a single operation.
Inside a GPU
- Streaming multiprocessors (SMs) are the main building blocks. Each one holds many CUDA cores, Tensor Cores, registers and fast shared memory.
- Tensor Cores are units built to multiply and accumulate small matrices very quickly. They are the main source of AI speed on modern GPUs.
- Precision formats: AI work often uses lower-precision numbers such as FP16, bfloat16 and TF32 to go faster and use less memory. Newer GPU generations add even lower precisions such as FP8.
- Memory hierarchy: registers, then shared memory and caches, then the GPU's main memory. Data close to the cores is fastest.
- HBM (high-bandwidth memory) is stacked memory placed right beside the GPU chip. It gives far higher bandwidth than ordinary graphics memory, which matters because many AI workloads are limited by memory speed.
Systems: DGX and HGX
- DGX is NVIDIA's complete, ready-built AI system.
- HGX is the GPU baseboard that server makers build into their own systems.
Scale-up and scale-out
| Scale-up (inside a server) | Scale-out (between servers) | |
|---|---|---|
| Technology | NVLink and NVSwitch | InfiniBand, or Spectrum-X Ethernet |
| Purpose | Very fast GPU-to-GPU links, so GPUs share data far faster than over PCIe | Connect many servers into one cluster with low latency |
- NVSwitch lets all the GPUs in a system talk to each other at full speed.
- RDMA lets one machine read another's memory without involving the CPU, which keeps network latency low.
- DPUs (such as BlueField) take over networking, storage and security tasks from the CPU.
Storage, power and cooling
- Storage and data pipelines: training needs to feed data to the GPUs fast enough to keep them busy. Parallel file systems and fast networks help, and a slow pipeline leaves expensive GPUs idle.
- Power and cooling: GPU servers draw a lot of power per rack and make a lot of heat, so racks hold fewer servers than usual and often use liquid cooling.
Self-check
Why is a GPU better than a CPU for deep learning?
Deep learning is mostly large matrix math that can be split into thousands of parallel operations. A GPU has thousands of simpler cores, Tensor Cores built for that math, and very high memory bandwidth. A CPU has a few strong cores tuned for low latency on varied work.
What is the difference between NVLink and InfiniBand?
NVLink is the fast GPU-to-GPU link inside a server (scale-up). InfiniBand, or Spectrum-X Ethernet, connects servers to each other in a cluster (scale-out).
What is the difference between DGX and HGX?
DGX is NVIDIA's complete, ready-built system. HGX is the GPU baseboard that server makers build into their own systems.
More content for this area is coming, including a lab on precision formats and Tensor Cores.