NCA-AIIO study notes
Area 2

AI infrastructure

The hardware that AI runs on: GPUs, the memory and links around them, and the servers, networks and facilities that hold them.

CPU vs GPU

CPUGPU
CoresFew, strong cores (tens)Thousands of simpler cores
Built forLow latency on varied, branching workHigh throughput on the same operation applied to lots of data
Best atRunning the operating system, control logic, serial codeMatrix math, which is most of deep learning

Deep learning is mostly large matrix multiplications that split into many independent pieces, which is why GPUs fit it so well. See the labs for how a CPU core handles a single operation.

Inside a GPU

Systems: DGX and HGX

Scale-up and scale-out

Scale-up (inside a server)Scale-out (between servers)
TechnologyNVLink and NVSwitchInfiniBand, or Spectrum-X Ethernet
PurposeVery fast GPU-to-GPU links, so GPUs share data far faster than over PCIeConnect many servers into one cluster with low latency

Storage, power and cooling

Self-check

Why is a GPU better than a CPU for deep learning?

Deep learning is mostly large matrix math that can be split into thousands of parallel operations. A GPU has thousands of simpler cores, Tensor Cores built for that math, and very high memory bandwidth. A CPU has a few strong cores tuned for low latency on varied work.

What is the difference between NVLink and InfiniBand?

NVLink is the fast GPU-to-GPU link inside a server (scale-up). InfiniBand, or Spectrum-X Ethernet, connects servers to each other in a cluster (scale-out).

What is the difference between DGX and HGX?

DGX is NVIDIA's complete, ready-built system. HGX is the GPU baseboard that server makers build into their own systems.

More content for this area is coming, including a lab on precision formats and Tensor Cores.