ZICQ
中 Log in / Sign up
Newsroom Research & Papers #RAG #NLP #arXiv #Retrieval-Augmented Generation #Model Evaluation

arXiv Publishes Comprehensive Evaluation Framework for Retrieval-Augmented Generation Components

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:A new study published on arXiv presents a comprehensive evaluation framework for Retrieval-Augmented Generation (RAG) components, including 3 parsers, 3 chunking strategies, and 5 dense embedding models, along with a sparse BM25 baseline. The research, conducted across 800 question instances from four distinct Indian government regulatory documents, reveals that no single retriever family dominates across all documents, and the choice of parser and chunker significantly impacts performance. The


Background and Motivation

Retrieval-Augmented Generation (RAG) systems are typically assembled from independently-chosen components, including document parsers, chunking strategies, and embedding models. However, these choices are rarely evaluated jointly, and existing evaluations usually focus on a single document or corpus, making it difficult to understand the full impact of different component combinations on retrieval performance.

Methodology

The research team conducted a controlled factorial study evaluating 3 parsers, 3 chunking strategies, and 5 dense embedding models, along with a sparse BM25 baseline. The evaluation was performed on 800 question instances, each with one or more evidence strings, which were automatically validated against the source text and manually reviewed for a 10% random sample. The data covered four structurally distinct Indian central-government regulatory documents.

Key Findings

  1. No Single Dominant Retriever: No single retriever family dominates across all documents, and the choice of parser and chunker significantly impacts performance.
  2. MPNet-base Underperforms: MPNet-base shows a severe failure mode in handling table-derived questions.
  3. Near-Saturated Evidence Preservation: The corpus exhibits a near-saturated evidence-preservation ceiling above 98%, indicating that retrieval differences are driven primarily by ranking quality rather than information loss during ingestion.
  4. Impact of Embedding Dimensions and Chunk Sizes: The study also explores embedding-dimension and chunk-size/overlap ablations and conducts an efficiency/quality Pareto analysis.

Technical Highlights

  • Comprehensive Evaluation Framework: The study provides the first systematic evaluation of RAG component interactions, offering new insights for optimizing RAG systems.
  • Large-Scale Dataset: The 800-question benchmark dataset, covering diverse document structures, ensures a comprehensive and reliable evaluation.
  • Open-Source Toolkit: The research team releases the full evaluation harness and benchmark dataset, providing valuable resources for future research.

Industry Impact and Developer Recommendations

This study offers valuable insights for optimizing RAG systems, suggesting that developers should consider the interactions between parsers and chunking strategies and pay attention to the performance bottlenecks of embedding models. Additionally, the findings indicate that model selection and configuration need to be more cautious when dealing with complex documents.


Source: ArXiv NLP/LLM (cs.CL) (2026-09-29)

— END —

Tags: #RAG #NLP #arXiv #Retrieval-Augmented Generation #Model Evaluation

Community Comments

Loading live comments and annotations…