AI operations
Running AI hardware day to day: scheduling work, managing GPU software, monitoring health and fixing problems.
Orchestration: Slurm and Kubernetes
| Slurm | Kubernetes | |
|---|---|---|
| Typical use | Batch jobs on HPC and training clusters | Containerized services and pipelines, including inference |
| How work is submitted | A job script that asks for nodes, GPUs and time | A manifest describing containers and the resources they need |
| Scheduling | A queue with priorities and reservations | The scheduler places pods on nodes with free resources |
GPU software on Kubernetes
- The NVIDIA GPU Operator automates installing and managing the GPU software on Kubernetes nodes: the driver, the container toolkit, the device plugin and monitoring components.
- Containers and NGC: applications and frameworks ship as containers, and NVIDIA's NGC catalog provides ready-made, GPU-optimized ones.
Sharing a GPU: MIG and vGPU
- MIG (Multi-Instance GPU) splits one physical GPU into several isolated hardware partitions, each with its own memory and compute. It is only available on supported GPUs.
- vGPU shares a GPU between virtual machines through virtualization software.
Monitoring and troubleshooting
nvidia-smigives a quick look at one machine's GPUs: utilization, memory, temperature and processes.- DCGM (Data Center GPU Manager) monitors and checks the health of GPUs across a whole fleet, and can feed metrics to tools such as Prometheus and Grafana.
- Common problems to recognize: thermal throttling, memory (ECC) errors, GPU error codes reported as Xid messages, a driver or container mismatch, and jobs that cannot get GPUs because they are all in use.
Planning and measuring
- Benchmarking shows whether a cluster delivers the performance it should, for example how well training scales as GPUs are added.
- Capacity planning matches expected workloads to GPUs, network, storage, power and cooling before demand outgrows them.
Self-check
What does the GPU Operator do?
It automates installing and managing the GPU software on Kubernetes nodes: the driver, container toolkit, device plugin and monitoring components.
What is the difference between MIG and vGPU?
MIG splits one physical GPU into several isolated hardware partitions with their own memory and compute. vGPU shares a GPU between virtual machines using virtualization software.
What would you use for a quick check of one machine's GPUs, and what for a whole fleet?
nvidia-smi for a quick look at one machine. DCGM for monitoring and health checks across many GPUs.
More content for this area is coming, including a lab on how a Slurm job moves through the queue.