ZICQ
中 Log in / Sign up
Newsroom Research & Papers #Hugging Face #Split Learning #Text Leakage #Privacy Protection #Gradient Analysis

Hugging Face Research Reveals Text Leakage and Gradient Impact in Split Learning

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face's research team has published a study on text leakage mechanisms in split learning, highlighting the significant impact of gradients on text reconstruction during model training. The study shows that an attacker can recover 94.20% of tokens from activations alone using only the publicly released weights of the client layers, and this increases to 97.38% when gradients are also available. Additionally, the research examines leakage across different document lengths and proposes evalu


Background and Motivation

Split learning is an emerging privacy-preserving machine learning approach that allows a client to train a language model on a server without sending its text data. The client runs the initial layers of the model and sends the output (a numerical vector for each token) to the server. During training, the server sends gradients back to the client. However, this method may introduce risks of text leakage.

Key Findings

  1. Impact of Gradients on Text Reconstruction

    • The study shows that an attacker can recover 94.20% of tokens from activations alone using only the publicly released weights of the client layers.
    • When gradients are also available, the recovery rate increases to 97.38%, a 3.17 percentage point increase.
  2. Document-Level Leakage

    • For 32-token documents, the attacker can reconstruct 13.71% of documents exactly without gradients, and 37.77% with gradients.
    • This difference indicates that gradients have a more significant impact on the reconstruction of longer texts.
  3. Evaluation of Defense Strategies

    • The study tested the 'Secret Mixup' technique, which blends each outgoing vector with a decoy to prevent attackers from reconstructing documents exactly.
    • Although this technique almost prevents the exact reconstruction of any document, the attacker can still recover 83-91% of tokens, showing the limitations of current defense strategies.
  4. Influence of Model Training Layers

    • Experiments on GPT-2 and Qwen3-0.6B models show that the starting position of the server-trained layers affects both model quality and leakage, even when the length of the training layers is fixed.

Recommendations and Conclusions

  • Leakage Evaluation Method: It is recommended to report leakage on both a per-token and per-document basis.
  • Data Sensitivity: Data sent by split models should be treated as sensitive as the text itself.
  • Defense Strategy Improvement: Current defense strategies, such as Secret Mixup, can prevent exact reconstruction of documents but cannot completely prevent token-level leakage. Further research on more effective defense mechanisms is needed.

Industry Impact

This research poses new challenges to the security of split learning and emphasizes the importance of data privacy protection. The findings suggest that when designing and using split learning methods, the risk of data leakage must be fully considered, and stricter protection measures should be taken.


Source: Hugging Face Daily Papers (2026-10-02)

— END —

Tags: #Hugging Face #Split Learning #Text Leakage #Privacy Protection #Gradient Analysis

Community Comments

Loading live comments and annotations…