Skip to content
BITBRIEF

Institutional research · AI · Cybersecurity · Digital assets

Vol. 01 · No. 13

Chunking

Also: Document splitting · Passage segmentation

Splitting documents into passages before they are embedded and indexed — the step that decides what a retrieval system is able to find at all.

A retrieval system searches passages, not documents. Chunking is where the passages get made, and a fact split across a boundary is a fact the system cannot retrieve intact no matter how good the embedding model or the ranker downstream.

Why it survived larger context windows

Bigger windows changed how much can be sent, not what gets found. If the chunk that contained the answer was never retrieved, the size of the window is irrelevant — and if everything is sent instead, cost and latency rise while precision falls.

Where it goes wrong

  • Fixed-length splitting that cuts through tables, clauses and definitions, leaving both halves useless.
  • Chunks that carry no context of their own — a passage saying 'this must be reported within 30 days' is unusable when 'this' was in the previous chunk.
  • Re-chunking without re-embedding, which silently leaves an index describing a document that no longer exists in that form.

It is unglamorous work and it dominates outcomes. Retrieval quality is decided at index time far more than at query time.

All 30 terms