ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #MoE Architecture #Privacy Protection #Membership Inference Attack #arXiv

arXiv Research Reveals: Router Telemetry in MoE Models Vulnerable to Membership Inference Attacks

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:A new study on arXiv highlights that the routing telemetry data generated by Mixture-of-Experts (MoE) language models during inference can leak privacy-sensitive information about the model's training data. The researchers introduced a router-augmented membership inference attack that combines conventional output signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models. Across various architectures and data domains, the tele


Background and Problem

Mixture-of-Experts (MoE) language models generate routing information during inference to optimize computational efficiency, and this telemetry data is often logged for monitoring, debugging, load analysis, and safety auditing. However, this data can expose a view of the model's internal computation, raising a critical question: can it reveal whether a particular example was used to fine-tune the deployed model?

Methodology and Findings

The research team introduced a router-augmented membership inference attack with the following steps:

  1. Attack Method: Combine conventional output signals with aggregated routing features and apply a membership classifier learned from independently fine-tuned shadow models.
  2. Experimental Results: Across three MoE architectures and three data domains, the telemetry consistently improved the attack's True Positive Rate (TPR) by 2.7 to 9.4 percentage points at a 1% False Positive Rate (FPR).
  3. Wide Applicability: The leakage persists across full fine-tuning, frozen-router training, LoRA, and instruction tuning. Even with discrete expert selections, restricted telemetry, or a single shadow model, the attack remains effective.
  4. Mechanistic Analysis: The study shows that the leakage does not require router-specific memorization. Fine-tuning introduces membership information into hidden representations, while the router exposes a projection of this signal even when its parameters are frozen.

Technical Implications and Recommendations

  1. Privacy Risks: This research underscores the potential privacy risks associated with MoE models, urging developers to exercise caution when using MoE architectures.
  2. Improvement Directions: Future research should focus on developing stronger privacy-preserving mechanisms, such as differential privacy techniques or hardware-based solutions, to mitigate the privacy leakage risks of routing telemetry data.
  3. Developer Recommendations: When using MoE models, developers should minimize the logging and exposure of routing telemetry data and consider introducing privacy-preserving technologies during model training and inference.

Conclusion

This study provides a new perspective on the privacy protection of MoE models, emphasizing the importance of routing telemetry data in membership inference attacks. Future research can further explore more effective privacy protection methods to ensure the security of MoE models in practical applications.


Source: ArXiv Machine Learning (cs.LG) (2026-10-09)

— END —

Tags: #MoE Architecture #Privacy Protection #Membership Inference Attack #arXiv

Community Comments

Loading live comments and annotations…