ZICQ
中 Log in / Sign up
Newsroom Research & Papers #Hugging Face #Multi-Vector Retrieval #Data Security #Reverse Engineering #AI Safety

Hugging Face Unveils Inversion Vulnerability in Multi-Vector Visual Document Retrievers: Indices Can Be Reconstructed

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face researchers have discovered a critical security vulnerability in multi-vector visual document retrievers. Despite storing documents as thousands of vectors in third-party databases, the team successfully reconstructed document content from the indices alone by exploiting the regularity of the stored vectors. In the ViDoRe v3 benchmark, inverted pages recovered 47% of the words and 45% of the sensitive tokens, and queries against stored indices ranked the source page first 98.4% of t


Background and Problem

Multi-vector visual document retrievers divide documents into thousands of vectors and store them in vector databases, providing efficient retrieval capabilities. However, the security implications of this approach have been largely overlooked. Hugging Face researchers discovered that these indices are not as secure as previously thought and can be used to reconstruct the original document content.

Methodology and Findings

The researchers framed the inversion problem as a conditional document image generation task and inferred the critical information needed for the attack: the encoder, the page shape, and the order of the vectors. In the ViDoRe v3 benchmark, inverted pages recovered 47% of the words and 45% of the sensitive tokens, and queries against stored indices ranked the source page first 98.4% of the time.

Protection Measures and Effectiveness

The study tested two inexpensive protection measures:

  • Token Pooling: Merging vectors of multiple tokens into one to reduce information.
  • Shuffling: Randomly shuffling the order of the vectors. Both methods reduced word recall to about 8%. However, the researchers also found that training a model to restore the order of a shuffled index could increase the matching rate of source pages from 3.8% to 93.5%. This indicates that current protection measures have significant limitations.

Cross-Model Validation and Generalization

To verify the universality of the vulnerability, the researchers applied the same attack method to another multi-vector retriever. The results showed that inverted pages still ranked their source page first 70.2% of the time, although the word recall remained below a nearest-neighbor baseline.

Conclusion and Recommendations

Multi-vector visual document retrievers have significant security vulnerabilities in their index storage. The study recommends treating indices as equally sensitive as the documents they encode and implementing stronger protection measures.

Industry Impact and Developer Recommendations

  • Data Privacy Protection: Developers should reassess the security of multi-vector retrievers and consider adopting more advanced encryption or obfuscation techniques.
  • Index Protection Strategies: It is recommended to encrypt indices during storage and transmission and to regularly update index structures to increase inversion difficulty.
  • Future Research Directions: Explore more effective protection measures, such as AI-based dynamic index protection mechanisms.

Source: Hugging Face Daily Papers (2026-10-07)

— END —

Tags: #Hugging Face #Multi-Vector Retrieval #Data Security #Reverse Engineering #AI Safety

Community Comments

Loading live comments and annotations…