Document Chunking
Document chunking divides source material into retrievable, attributable passages for embedding or generation. Each passage needs useful context while fitting downstream input limits.
Common approaches
| Approach | Idea | Check |
|---|---|---|
| Fixed length | Character or token windows | Broken sentences and structures |
| Sentence or heading | Respect natural boundaries | Uneven passage sizes |
| Structure-aware | Parse Markdown, HTML, or code | Table, code, and heading integrity |
| Semantic | Use changes in neighboring meaning | Whether added computation helps retrieval |
LlamaIndex provides node parsers using these approaches. Evaluate the choice on your own questions.
A proposed experiment
Try 256-, 512-, and 1024-token windows as illustrative candidates. Compare whether evidence stays complete, retrieval duplicates passages, and the assembled context fits the generator. These are trial settings, not recommended optima.
Overlap preserves boundary context at the cost of additional stored text and repeated candidates. State whether length is measured in characters or tokens.
Preserve attribution
We suggest retaining document and chunk IDs, heading paths, source locations, versions, and access metadata. An instruction should lead back to the corresponding section of the original help page.
Inspect failure cases
Our suggested manual review includes fragmented tables, missing headings, sentences broken across pages, and duplicate passages consuming the context budget. Similarity scores alone do not establish a good split.
Sources
- LlamaIndex: Node Parser Modules — strategies and parameters.
- LlamaIndex: Documents / Nodes — nodes, metadata, and source relationships.