ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #RBS-Attention #Long-Context #Sparse Prefill #LLMs & Foundation Models #Inference Efficiency

arXiv Introduces RBS-Attention: Revolutionizing Sparse Prefill for Long-Context LLMs

Avatar of Mr.Xu

By Mr.Xu

Published: · 6 views

中文阅读 (Chinese) English Version

Summary:arXiv has introduced RBS-Attention, a novel training-free sparse-prefill method for long-context large language models (LLMs). This approach addresses the 'mean dilution' problem in traditional sparse prefill by employing a dual-branch selection mechanism: a centroid base branch for average relevance and a rescue branch for identifying underestimated risks using maximum key-block radius. On H100 GPUs, RBS-Attention achieves a 20.65x standalone prefill-attention speedup and a 5.97x end-to-end tim


1. Background and Motivation

The inference process of long-context large language models (LLMs) is limited by the prefill stage, where dense self-attention processes the entire prompt, leading to inefficiency. Sparse block selection techniques can reduce computational costs, but centroid-based methods may hide highly relevant tokens, causing the 'mean dilution' problem.

2. Overview of RBS-Attention

RBS-Attention is a training-free sparse-prefill method that employs a dual-branch selection mechanism:

  • Centroid Base Branch: Captures average relevance.
  • Rescue Branch: Uses the maximum key-block radius and its distribution related to the prompt, layer, and head to identify blocks at risk of underestimation.

By independently thresholding the two branches and combining their masks, RBS-Attention controls the contribution of rescue blocks while maintaining regular block-sparse FlashAttention execution.

3. Experimental Results

  • Speedup: On H100 GPUs, RBS-Attention achieves a 20.65x standalone prefill-attention speedup, 11.92x vLLM prefill-attention speedup, and a 5.97x end-to-end time-to-first-token speedup at 128K context length on Qwen3-30B-A3B-Instruct-2507-FP8.
  • Accuracy: On the dense Qwen3-32B model, RBS-Attention obtains 88.65 overall RULER accuracy compared to 89.52 for dense attention. LongBench-v2, InfiniteBench, and Video-MME provide additional quality evaluation.

4. Technical Highlights

  • Training-Free Method: Achieves efficient sparse prefill without additional training.
  • Dual-Branch Mechanism: Combines centroid and rescue branches to effectively address the 'mean dilution' problem.
  • Hardware Acceleration: Achieves significant speedup on H100 GPUs.

5. Industry Impact and Developer Recommendations

RBS-Attention offers a new approach to improving the inference efficiency of long-context LLMs, particularly for applications requiring processing extremely long texts, such as legal document analysis, academic paper summarization, and cross-modal reasoning. Developers can integrate this method into existing LLM inference pipelines to enhance overall performance.


Source: ArXiv AI (cs.AI) (2026-09-21)

— END —

Tags: #RBS-Attention #Long-Context #Sparse Prefill #LLMs & Foundation Models #Inference Efficiency

Community Comments

Loading live comments and annotations…