← all topics

Retrieval Augmented Generation

RAG

Give the model the right documents at query time instead of hoping it memorised them.

Levels foundation / engineer / advanced / architect
Depth 9
Time 5h
Kind system
On AI Engineer, RAG Specialist, AI Architect

Grasp

A language model knows what was in its training data, up to a cutoff, with no citations and no awareness of your organisation. Retrieval-augmented generation fixes that by retrieving relevant documents at query time and placing them in the context window before the model answers.

The loop is short: embed the query, search a corpus for the closest passages, insert them into the prompt, generate an answer grounded in them. Four steps, and the naive version can be built in an afternoon.

Which is exactly the trap. The demo works immediately and the production system does not, because quality is dominated by retrieval rather than generation. If the right passage is not in the top results, no model can save the answer. Nearly every child node below this one - chunking, hybrid search, reranking, evaluation - exists to address a specific way that retrieval quietly fails.