Hugging Face Study: Mixed Supervised Fine-Tuning Outperforms Next-Chunk Reasoning RL on No-CoT Data Training
By Mr.Xu
Published:
Summary:Hugging Face has released a study comparing Next-Chunk Reasoning RL, which leverages no-CoT (no Chain-of-Thought) data, with a simpler alternative called Mixed SFT (Supervised Fine-Tuning). Mixed SFT, which jointly trains on no-CoT and long-CoT data, achieves a higher performance ceiling in post-RLVR tasks while requiring over 60 times less training compute. This approach demonstrates consistent advantages across in-domain mathematical reasoning and out-of-domain tasks, challenging the dominance
Background and Motivation
In recent years, reinforcement learning (RL) has been widely applied in natural language processing for reasoning tasks, especially when dealing with no-CoT (no Chain-of-Thought) data. No-CoT data contains rich reasoning content but lacks explicit reasoning chain annotations. Existing research has proposed Next-Chunk Reasoning RL, which trains a model to generate implicit reasoning traces and rewards them based on their ability to predict the next chunk of text. However, current evaluations primarily compare against traditional supervised fine-tuning (SFT) baselines, leaving open whether the gains come from the RL formulation itself or from more effectively leveraging no-CoT data.
Methodology and Findings
Hugging Face's research team conducted a controlled study comparing Next-Chunk Reasoning RL with a novel alternative called Mixed SFT (Supervised Fine-Tuning). Mixed SFT is a simple single-stage supervised fine-tuning method that jointly trains on no-CoT and long-CoT data. Despite its simplicity, Mixed SFT achieves a significantly higher performance ceiling in post-RLVR tasks while requiring over 60 times less training compute. This advantage is consistent across both in-domain mathematical reasoning and out-of-domain tasks.
Furthermore, the study found that higher pre-RLVR accuracy does not necessarily translate into higher post-RLVR accuracy. This highlights the importance of evaluating no-CoT training strategies within the full training pipeline.
Technical Highlights
- Mixed Supervised Fine-Tuning (Mixed SFT): A simple yet effective training method that jointly trains on no-CoT and long-CoT data, significantly boosting model performance.
- Training Efficiency: Mixed SFT offers a high-performance solution with substantially reduced training compute.
- Cross-Domain Applicability: The method demonstrates strong performance in both mathematical reasoning and cross-domain tasks, showcasing its broad applicability.
Industry Impact and Developer Recommendations
This research holds significant implications for AI developers. It challenges the dominance of RL-based methods in no-CoT data training and provides new insights into more efficient training strategies. The efficiency of Mixed SFT makes it an ideal choice for resource-constrained environments.
For developers, it is recommended to consider mixed supervised fine-tuning methods when designing no-CoT data training strategies and to evaluate them within the full training pipeline. Additionally, developers should focus on balancing training efficiency with model performance to achieve optimal resource utilization and performance.
— END —Source: Hugging Face Daily Papers (2026-08-24)
Tags: #Hugging Face #Reinforcement Learning #Supervised Fine-Tuning #Reasoning Tasks #Training Efficiency
Community Comments