ZICQ
中 Log in / Sign up
ZICQ Info LLMs & Foundation Models #ArXiv #Vision-Language Models #Pruning Technique #Inference Efficiency #Dempster-Shafer Theory

ArXiv Proposes E2S-Pruner: Revolutionizing Visual Token Pruning in Vision-Language Models

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:The ArXiv team introduces E2S-Pruner, a novel method for visual token pruning in vision-language models. This progressive two-stage evidence fusion framework achieves efficient pruning without requiring auxiliary models, trainable parameters, or fine-tuning. By estimating the reliability of attention heads and using Dempster-Shafer evidence theory to fuse multi-layer evidence, E2S-Pruner significantly improves model throughput while maintaining high performance across multiple benchmarks.


Background and Challenges

Vision-language models (VLMs) typically encode images into hundreds of visual tokens, which introduces significant inference latency and GPU memory overhead. Existing pruning methods rely heavily on attention scores and directly aggregate outputs across attention heads and network layers, making it difficult to characterize evidential uncertainty and conflict.

E2S-Pruner Method

The ArXiv team addresses these issues with E2S-Pruner through the following steps:

  1. Stage 1: Evidence Reliability Estimation

    • Treat each attention head as an independent evidence source.
    • Estimate its reliability based on evidence clarity and inter-head consistency.
    • Represent each visual token using three states: important, unimportant, and uncertain.
  2. Stage 2: Evidence Fusion

    • Use Dempster-Shafer evidence theory to quantify inter-layer conflict.
    • Fuse complementary evidence from multiple network layers.
    • Introduce a spatial novelty constraint to promote coverage of distinct image regions and prevent retained tokens from concentrating in a few locally salient areas.

Experimental Results

On the LLaVA-1.5-7B model, E2S-Pruner retains 98.0%, 96.8%, and 90.6% of the aggregate performance when the average numbers of retained visual tokens are 192, 128, and 64, respectively, while improving throughput by 1.96x and 2.09x under the 128-token and 64-token settings. Experiments on Qwen2-VL-7B further demonstrate the cross-model generalization capability of the method.

Technical Highlights

  • No Auxiliary Models or Parameters: E2S-Pruner achieves efficient pruning without requiring auxiliary models or trainable parameters.
  • Evidence Fusion and Conflict Quantification: Utilizes Dempster-Shafer theory to fuse multi-level evidence and quantify inter-layer conflict.
  • Spatial Novelty Constraint: Introduces a spatial novelty constraint to ensure retained tokens cover different image regions.
  • Cross-Model Generalization: Experiments on multiple models demonstrate the method's versatility.

Industry Impact and Developer Recommendations

E2S-Pruner offers a new approach to improving the inference efficiency of vision-language models, particularly in resource-constrained environments. Developers can consider the following recommendations:

  • Model Optimization: When deploying vision-language models on resource-constrained devices, consider using E2S-Pruner for pruning.
  • Performance Evaluation: Conduct comprehensive performance evaluations after applying E2S-Pruner to ensure the pruned model maintains sufficient accuracy for critical tasks.
  • Cross-Model Application: Explore applying the method to other types of models to explore its broader application potential.

Source: ArXiv cs.AI (2026-08-24)

— END —

Tags: #ArXiv #Vision-Language Models #Pruning Technique #Inference Efficiency #Dempster-Shafer Theory

Community Comments

Loading live comments and annotations…