← all topics

Retrieval Augmented Generation

Chunking

How you split documents. Where most RAG systems quietly lose their quality.

Levels engineer / advanced
Depth 10
Time 5h
Kind technique
On RAG Specialist

Grasp

Retrieval works on units, and chunking is where those units are decided. A document has to be cut into pieces small enough that a piece is mostly about one thing, because an embedding averages whatever it is given and a long passage covering four topics embeds as a vague blur that matches nothing well. Cut too small and you win precision but lose the sentence that made the fact meaningful.

Most of the damage is done by cutting blindly. Splitting on a fixed number of characters ends a chunk mid-clause, separates a table from its header, and detaches a figure's caption from the figure. Splitting on the document's own structure instead, by heading, section, paragraph or list item, keeps pieces that were written as a unit together. Overlap is the usual patch for boundary loss, and it works, but it is a tax: every chunk you duplicate is paid for again at embedding time, at storage, and on every scan.

The part that repays the most effort is context. A chunk retrieved on its own has lost the title it sat under, the document it came from and the date it was written, so the model receives a confident paragraph with no provenance. Carrying that heading path and metadata into the chunk's text is usually a bigger quality win than any tuning of its length.