Skip to main content
ZICQ

Wiki Terminology

Document Chunking

Terminology
Aliases: text splitting chunks ·2026-10-07

Document Chunking

Document chunking divides source material into retrievable, attributable passages for embedding or generation. Each passage needs useful context while fitting downstream input limits.

Common approaches

Approach Idea Check
Fixed length Character or token windows Broken sentences and structures
Sentence or heading Respect natural boundaries Uneven passage sizes
Structure-aware Parse Markdown, HTML, or code Table, code, and heading integrity
Semantic Use changes in neighboring meaning Whether added computation helps retrieval

LlamaIndex provides node parsers using these approaches. Evaluate the choice on your own questions.

A proposed experiment

Try 256-, 512-, and 1024-token windows as illustrative candidates. Compare whether evidence stays complete, retrieval duplicates passages, and the assembled context fits the generator. These are trial settings, not recommended optima.

Overlap preserves boundary context at the cost of additional stored text and repeated candidates. State whether length is measured in characters or tokens.

Preserve attribution

We suggest retaining document and chunk IDs, heading paths, source locations, versions, and access metadata. An instruction should lead back to the corresponding section of the original help page.

Inspect failure cases

Our suggested manual review includes fragmented tables, missing headings, sentences broken across pages, and duplicate passages consuming the context budget. Similarity scores alone do not establish a good split.

Sources