ArXiv Proposes E2S-Pruner: Revolutionizing Visual Token Pruning in Vision-Language Models
By Mr.Xu
Published:
Summary:The ArXiv team introduces E2S-Pruner, a novel method for visual token pruning in vision-language models. This progressive two-stage evidence fusion framework achieves efficient pruning without requiring auxiliary models, trainable parameters, or fine-tuning. By estimating the reliability of attention heads and using Dempster-Shafer evidence theory to fuse multi-layer evidence, E2S-Pruner significantly improves model throughput while maintaining high performance across multiple benchmarks.
Background and Challenges
Vision-language models (VLMs) typically encode images into hundreds of visual tokens, which introduces significant inference latency and GPU memory overhead. Existing pruning methods rely heavily on attention scores and directly aggregate outputs across attention heads and network layers, making it difficult to characterize evidential uncertainty and conflict.
E2S-Pruner Method
The ArXiv team addresses these issues with E2S-Pruner through the following steps:
-
Stage 1: Evidence Reliability Estimation
- Treat each attention head as an independent evidence source.
- Estimate its reliability based on evidence clarity and inter-head consistency.
- Represent each visual token using three states: important, unimportant, and uncertain.
-
Stage 2: Evidence Fusion
- Use Dempster-Shafer evidence theory to quantify inter-layer conflict.
- Fuse complementary evidence from multiple network layers.
- Introduce a spatial novelty constraint to promote coverage of distinct image regions and prevent retained tokens from concentrating in a few locally salient areas.
Experimental Results
On the LLaVA-1.5-7B model, E2S-Pruner retains 98.0%, 96.8%, and 90.6% of the aggregate performance when the average numbers of retained visual tokens are 192, 128, and 64, respectively, while improving throughput by 1.96x and 2.09x under the 128-token and 64-token settings. Experiments on Qwen2-VL-7B further demonstrate the cross-model generalization capability of the method.
Technical Highlights
- No Auxiliary Models or Parameters: E2S-Pruner achieves efficient pruning without requiring auxiliary models or trainable parameters.
- Evidence Fusion and Conflict Quantification: Utilizes Dempster-Shafer theory to fuse multi-level evidence and quantify inter-layer conflict.
- Spatial Novelty Constraint: Introduces a spatial novelty constraint to ensure retained tokens cover different image regions.
- Cross-Model Generalization: Experiments on multiple models demonstrate the method's versatility.
Industry Impact and Developer Recommendations
E2S-Pruner offers a new approach to improving the inference efficiency of vision-language models, particularly in resource-constrained environments. Developers can consider the following recommendations:
- Model Optimization: When deploying vision-language models on resource-constrained devices, consider using E2S-Pruner for pruning.
- Performance Evaluation: Conduct comprehensive performance evaluations after applying E2S-Pruner to ensure the pruned model maintains sufficient accuracy for critical tasks.
- Cross-Model Application: Explore applying the method to other types of models to explore its broader application potential.
— END —Source: ArXiv cs.AI (2026-08-24)
Tags: #ArXiv #Vision-Language Models #Pruning Technique #Inference Efficiency #Dempster-Shafer Theory
Community Comments