arXiv Introduces RBS-Attention: Revolutionizing Sparse Prefill for Long-Context LLMs
By Mr.Xu
Published: · 6 views
Summary:arXiv has introduced RBS-Attention, a novel training-free sparse-prefill method for long-context large language models (LLMs). This approach addresses the 'mean dilution' problem in traditional sparse prefill by employing a dual-branch selection mechanism: a centroid base branch for average relevance and a rescue branch for identifying underestimated risks using maximum key-block radius. On H100 GPUs, RBS-Attention achieves a 20.65x standalone prefill-attention speedup and a 5.97x end-to-end tim
1. Background and Motivation
The inference process of long-context large language models (LLMs) is limited by the prefill stage, where dense self-attention processes the entire prompt, leading to inefficiency. Sparse block selection techniques can reduce computational costs, but centroid-based methods may hide highly relevant tokens, causing the 'mean dilution' problem.
2. Overview of RBS-Attention
RBS-Attention is a training-free sparse-prefill method that employs a dual-branch selection mechanism:
- Centroid Base Branch: Captures average relevance.
- Rescue Branch: Uses the maximum key-block radius and its distribution related to the prompt, layer, and head to identify blocks at risk of underestimation.
By independently thresholding the two branches and combining their masks, RBS-Attention controls the contribution of rescue blocks while maintaining regular block-sparse FlashAttention execution.
3. Experimental Results
- Speedup: On H100 GPUs, RBS-Attention achieves a 20.65x standalone prefill-attention speedup, 11.92x vLLM prefill-attention speedup, and a 5.97x end-to-end time-to-first-token speedup at 128K context length on Qwen3-30B-A3B-Instruct-2507-FP8.
- Accuracy: On the dense Qwen3-32B model, RBS-Attention obtains 88.65 overall RULER accuracy compared to 89.52 for dense attention. LongBench-v2, InfiniteBench, and Video-MME provide additional quality evaluation.
4. Technical Highlights
- Training-Free Method: Achieves efficient sparse prefill without additional training.
- Dual-Branch Mechanism: Combines centroid and rescue branches to effectively address the 'mean dilution' problem.
- Hardware Acceleration: Achieves significant speedup on H100 GPUs.
5. Industry Impact and Developer Recommendations
RBS-Attention offers a new approach to improving the inference efficiency of long-context LLMs, particularly for applications requiring processing extremely long texts, such as legal document analysis, academic paper summarization, and cross-modal reasoning. Developers can integrate this method into existing LLM inference pipelines to enhance overall performance.
— END —Source: ArXiv AI (cs.AI) (2026-09-21)
Tags: #RBS-Attention #Long-Context #Sparse Prefill #LLMs & Foundation Models #Inference Efficiency
Community Comments