NCA-AIIO study notes
Area 3

AI operations

Running AI hardware day to day: scheduling work, managing GPU software, monitoring health and fixing problems.

Orchestration: Slurm and Kubernetes

SlurmKubernetes
Typical useBatch jobs on HPC and training clustersContainerized services and pipelines, including inference
How work is submittedA job script that asks for nodes, GPUs and timeA manifest describing containers and the resources they need
SchedulingA queue with priorities and reservationsThe scheduler places pods on nodes with free resources

GPU software on Kubernetes

Sharing a GPU: MIG and vGPU

Monitoring and troubleshooting

Planning and measuring

Self-check

What does the GPU Operator do?

It automates installing and managing the GPU software on Kubernetes nodes: the driver, container toolkit, device plugin and monitoring components.

What is the difference between MIG and vGPU?

MIG splits one physical GPU into several isolated hardware partitions with their own memory and compute. vGPU shares a GPU between virtual machines using virtualization software.

What would you use for a quick check of one machine's GPUs, and what for a whole fleet?

nvidia-smi for a quick look at one machine. DCGM for monitoring and health checks across many GPUs.

More content for this area is coming, including a lab on how a Slurm job moves through the queue.