NCA-AIIO study notes
Area 1

Essential AI knowledge

The vocabulary and the big picture: what AI workloads do, and which NVIDIA software supports each stage.

AI, machine learning and deep learning

Training vs inference

TrainingInference
What it doesAdjusts the model's weights using large amounts of dataUses the trained model to make predictions on new input
ComputeVery heavy, often many GPUs working togetherLighter per request, but runs constantly
What mattersTotal time to train, and scaling across GPUsLatency (how fast one answer) and throughput (answers per second)

The AI project lifecycle

  1. Collect and prepare data.
  2. Build and train the model.
  3. Validate it on data it has not seen.
  4. Deploy it for inference.
  5. Monitor it, and retrain when its quality drops.

Generative AI, LLMs and RAG

The NVIDIA software stack

PieceWhat it is for
CUDAThe platform and programming model for running general-purpose code on NVIDIA GPUs
cuDNNGPU-optimized building blocks for deep learning, used by frameworks such as PyTorch
TensorRTOptimizes a trained model so it runs inference faster
Triton Inference ServerServes models to applications at scale
NIMPrebuilt inference microservices that package optimized models for easy deployment
NeMoTools for building, customizing and training generative AI models
RAPIDSGPU-accelerated data science and data processing libraries
NVIDIA AI EnterpriseA supported software platform for running AI in production
NGCNVIDIA's catalog of containers, models and other software

Self-check

What is the difference between training and inference?

Training adjusts a model's weights using large amounts of data and is compute-heavy, often across many GPUs. Inference uses the trained model on new input, where latency and throughput matter most.

What problem does RAG address?

It lets an LLM answer from current or private documents by retrieving relevant passages and adding them to the prompt, instead of relying only on what it learned during training.

Which NVIDIA tool optimizes a trained model for inference, and which one serves it?

TensorRT optimizes the model. Triton Inference Server serves it.

More content for this area is coming. Check the topic list on NVIDIA's official page to see what to prioritize.