ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #RAG #Retrieval-Augmented Generation #NLP #Large Language Models #Multi-Document Evaluation

ArXiv Publishes Study on Retrieval-Augmented Generation for Indian Government Regulatory Documents

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:ArXiv has released a study on Retrieval-Augmented Generation (RAG) systems, focusing on the impact of document parsers, chunking strategies, and embedding models on RAG performance. The research conducted a controlled factorial study evaluating three parsers, three chunking strategies, and five dense embedding models, along with a sparse BM25 baseline, using 800 question instances across four distinct Indian government regulatory documents. The findings indicate that parser and chunker choices s


Background and Motivation

Retrieval-Augmented Generation (RAG) systems are typically assembled from independently-chosen components, including document parsers, chunking strategies, and embedding models. However, these choices are rarely evaluated jointly, and existing evaluations usually focus on a single document or corpus. This approach may lead to an incomplete understanding of the system's overall performance, especially when dealing with complex documents.

Methodology

The research team conducted a controlled factorial study evaluating the following components:

  • 3 document parsers: For parsing documents with different structures.
  • 3 chunking strategies: For splitting documents into smaller units.
  • 5 dense embedding models: For generating vector representations of text.

Additionally, a sparse BM25 baseline model was included for comparison. The experiment involved 800 question instances, each with one or more evidence strings that required validation. The evidence strings were confirmed through automatic validation and a 10% random manual review.

The team fitted a linear mixed-effects model with document-query-level random intercepts to the resulting 72,000-row result set and performed Holm-corrected paired comparisons, reporting clustered bootstrap confidence intervals for all 54 unique retriever configurations.

Key Findings

  1. No Single Retriever Dominates: Different document types lead to significant variations in retriever performance, with no single retriever family outperforming others across all documents.
  2. Significant Interaction Between Parser and Chunker: The choice of parser and chunker significantly impacts system performance.
  3. MPNet-base Model Underperforms: The MPNet-base model shows severe deficiencies in handling table-derived questions.
  4. Near-Saturated Evidence Preservation: The corpus exhibits a near-saturated evidence preservation ceiling above 98%, indicating that retrieval differences are primarily driven by ranking quality rather than information loss during ingestion.

Additional Analysis

The study also included ablation analyses on embedding dimensions and chunk size/overlap, as well as an efficiency/quality Pareto analysis. The full evaluation harness, corpus manifest, and 800-question benchmark were released.

Industry Impact and Developer Recommendations

This study provides new insights into the design and optimization of RAG systems:

  • Importance of Component Selection: The choice of parser, chunker, and embedding model should be optimized based on the specific application scenario.
  • Necessity of Multi-Document Evaluation: When evaluating RAG systems, it is crucial to consider different types of documents to obtain a comprehensive performance assessment.
  • Model Improvement Directions: Developers should focus on improving models for handling tabular data to enhance overall performance.

Conclusion

This research demonstrates the potential of RAG systems in handling complex documents and provides valuable references for future studies.


Source: ArXiv NLP/LLM (cs.CL) (2026-09-30)

— END —

Tags: #RAG #Retrieval-Augmented Generation #NLP #Large Language Models #Multi-Document Evaluation

Community Comments

Loading live comments and annotations…