ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #Policy Distillation #Cross-Tokenizer #Alignment Coverage #Supervision Reliability

Hugging Face Proposes New Cross-Tokenizer On-Policy Distillation Approach: Enhancing Alignment Coverage and Supervision

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face's research team has proposed an improved Cross-Tokenizer On-Policy Distillation (OPD) method that enhances learning efficiency by optimizing alignment coverage and supervision reliability. The approach, validated on mathematical reasoning and code generation tasks, demonstrates that restricting reverse KL divergence to a student-selected subset of the shared vocabulary at strictly aligned positions achieves accuracy comparable to full shared-vocabulary OPD while outperforming evalua


Background and Motivation

In the field of Natural Language Processing (NLP), On-Policy Distillation (OPD) is a method for training a student model using feedback from a teacher model. However, when the teacher and student models use different tokenizers, aligning their predictions becomes complex, requiring alignment at both the sequence and vocabulary levels. Hugging Face's research team has proposed an improved cross-tokenizer OPD approach that enhances learning efficiency by optimizing alignment coverage and supervision reliability.

Method and Experiments

The team conducted experiments on three heterogeneous teacher-student model pairs in mathematical reasoning and code generation tasks. The results show that strict 1:1 alignment groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the student model before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions.

By restricting the reverse KL divergence to a student-selected subset of the shared vocabulary (the top-16 at each strict position), the method achieves accuracy comparable to full shared-vocabulary OPD while outperforming evaluated baselines. However, adding span supervision reduces accuracy, indicating that prioritizing reliability over maximizing alignment coverage is more effective.

Key Findings and Contributions

  1. Alignment Coverage vs. Supervision Reliability: The study shows that prioritizing supervision reliability is more important than maximizing alignment coverage.
  2. Improvement in Cross-Tokenizer OPD: By restricting the reverse KL divergence to a subset of the shared vocabulary, efficient learning can be achieved.
  3. Limitations of Span Supervision: Adding span supervision reduces accuracy, suggesting the need for more refined control in the supervision process.

Industry Impact and Future Directions

This research provides a new technical pathway for cross-tokenizer OPD, particularly significant for handling model alignment issues with different tokenizers. In the future, this method could be applied in broader fields such as multilingual and multimodal models, further enhancing model learning efficiency and performance.

Developer Recommendations

For developers, this study offers a new perspective on handling cross-tokenizer model alignment issues by optimizing alignment coverage and supervision reliability to improve learning efficiency. Additionally, developers can experiment with applying this method to different tasks and models to verify its general applicability and effectiveness.


Source: Hugging Face Daily Papers (2026-10-06)

— END —

Tags: #Hugging Face #Policy Distillation #Cross-Tokenizer #Alignment Coverage #Supervision Reliability

Community Comments

Loading live comments and annotations…