ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #NavTree #Long-Document QA #Hierarchical Retrieval #arXiv #Retrieval-Augmented Generation

NavTree: A Novel Hierarchical Retrieval Method for Long-Document QA Outperforming Traditional Retrievors

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:A new study published on arXiv introduces NavTree, an innovative method for long-document question answering (QA). NavTree constructs a balanced segment tree and employs a hybrid lexical-and-dense frontier walk technique to achieve efficient navigation of lengthy documents. Compared to traditional flat retrievers and an extractive re-implementation of RAPTOR, NavTree demonstrates superior performance in matched-cost evaluations and significantly outperforms the BM25 baseline in long-document mul


Background and Motivation

In long-document question answering (QA) tasks, traditional flat retrieval methods like BM25 or top-$k$ retrieval often struggle to effectively capture complementary evidence within documents, leading to suboptimal performance. Methods like RAPTOR address this by constructing summary trees, but the process of generating summary content is complex and computationally expensive.

Overview of NavTree

NavTree introduces a novel hierarchical navigation approach that eliminates the need for language model-generated summaries. The core idea is to build a balanced segment tree and employ a hybrid lexical-and-dense frontier walk technique for navigation. The key steps are as follows:

  1. Building a Balanced Segment Tree: NavTree chunks the document during the indexing phase and constructs a deterministic balanced segment tree without relying on language models for summarization.
  2. Hybrid Frontier Walk: During the query phase, NavTree uses a hybrid lexical-and-dense frontier walk technique, starting from the root node and traversing the tree structure to emit only leaf node data for the reader.
  3. Navigation over Generation: NavTree emphasizes navigation over generating summary content, enhancing QA performance through efficient structured retrieval.

Experimental Results and Performance

NavTree demonstrates superior performance in long-document multi-hop QA tasks:

  • Outperforming BM25: NavTree is the only hierarchical method that significantly outperforms BM25 on a class-vs-class basis.
  • Matched-Cost Advantage: NavTree is the strongest matched-cost hierarchical retriever in the evaluated grid and ties the strongest flat baseline.
  • Zero Indexing Cost: Even at zero indexing cost, NavTree outperforms the extractive re-implementation of RAPTOR.

Furthermore, NavTree's performance remains consistent across stronger and open-weight readers and successfully isolates the structural advantage of leaf-only emission.

Technical Highlights

  • No Language Model-Generated Summaries: NavTree leverages structured navigation instead of generating summaries, significantly reducing computational costs.
  • Hybrid Lexical-and-Dense Frontier Walk: Combines the strengths of lexical matching and dense vector matching to enhance retrieval efficiency.
  • Balanced Segment Tree Structure: Ensures the stability and efficiency of the retrieval process through a deterministic balanced segment tree.

Industry Impact and Developer Recommendations

NavTree offers a new perspective for the long-document processing field, particularly in scenarios requiring efficient retrieval and QA, such as legal document analysis and medical literature retrieval. Developers can draw inspiration from NavTree's hierarchical navigation approach to optimize existing retrieval systems and improve long-document processing capabilities. Additionally, NavTree's zero-indexing cost characteristic makes it widely applicable in resource-constrained environments.


Source: ArXiv NLP/LLM (cs.CL) (2026-10-07)

— END —

Tags: #NavTree #Long-Document QA #Hierarchical Retrieval #arXiv #Retrieval-Augmented Generation

Community Comments

Loading live comments and annotations…