Essential AI knowledge
The vocabulary and the big picture: what AI workloads do, and which NVIDIA software supports each stage.
AI, machine learning and deep learning
- Artificial intelligence (AI) is the broad goal of software that does tasks needing human-like judgment.
- Machine learning (ML) is a way of doing AI where a model learns patterns from data instead of following hand-written rules.
- Deep learning (DL) is ML that uses neural networks with many layers. Most modern AI workloads on GPUs are deep learning.
Training vs inference
| Training | Inference | |
|---|---|---|
| What it does | Adjusts the model's weights using large amounts of data | Uses the trained model to make predictions on new input |
| Compute | Very heavy, often many GPUs working together | Lighter per request, but runs constantly |
| What matters | Total time to train, and scaling across GPUs | Latency (how fast one answer) and throughput (answers per second) |
The AI project lifecycle
- Collect and prepare data.
- Build and train the model.
- Validate it on data it has not seen.
- Deploy it for inference.
- Monitor it, and retrain when its quality drops.
Generative AI, LLMs and RAG
- Generative AI creates new content such as text, images or code.
- A large language model (LLM) is a generative model trained on huge amounts of text, which predicts the next token in a sequence.
- Retrieval-augmented generation (RAG) looks up relevant documents first and adds them to the prompt, so the model can answer from current or private information.
The NVIDIA software stack
| Piece | What it is for |
|---|---|
| CUDA | The platform and programming model for running general-purpose code on NVIDIA GPUs |
| cuDNN | GPU-optimized building blocks for deep learning, used by frameworks such as PyTorch |
| TensorRT | Optimizes a trained model so it runs inference faster |
| Triton Inference Server | Serves models to applications at scale |
| NIM | Prebuilt inference microservices that package optimized models for easy deployment |
| NeMo | Tools for building, customizing and training generative AI models |
| RAPIDS | GPU-accelerated data science and data processing libraries |
| NVIDIA AI Enterprise | A supported software platform for running AI in production |
| NGC | NVIDIA's catalog of containers, models and other software |
Self-check
What is the difference between training and inference?
Training adjusts a model's weights using large amounts of data and is compute-heavy, often across many GPUs. Inference uses the trained model on new input, where latency and throughput matter most.
What problem does RAG address?
It lets an LLM answer from current or private documents by retrieving relevant passages and adding them to the prompt, instead of relying only on what it learned during training.
Which NVIDIA tool optimizes a trained model for inference, and which one serves it?
TensorRT optimizes the model. Triton Inference Server serves it.
More content for this area is coming. Check the topic list on NVIDIA's official page to see what to prioritize.