Qwen Introduces Recursive Self-Rewrite (RSR) Framework: Enhancing Complex Task Solving
By Mr.Xu
Published:
Summary:The Qwen team introduces the Recursive Self-Rewrite (RSR) framework, a novel approach to improving model performance on complex tasks by leveraging diverse task-solving strategies. Using a base model, Qwen-3.8-27B, the framework discovers successful solutions under various constraints and reconstructs them as training trajectories for more efficient learning and performance gains. Experimental results demonstrate that the RSR framework significantly outperforms traditional methods, with pass@3 o
Recursive Self-Rewrite (RSR) Framework: Enhancing Complex Task Solving
The Qwen team introduces the Recursive Self-Rewrite (RSR) framework, a novel approach to overcoming model performance bottlenecks in complex tasks. The core idea of the RSR framework is to leverage diverse task-solving strategies to enhance the model's learning capabilities. The process involves the following steps:
-
Diverse Task-Solving Strategies: The RSR framework employs task-solving strategies under different constraints to generate diverse training trajectories. These strategies include using a planner to extract procedures, a critic to screen for verifier and solution leakage, and an executor to run qualified runbooks in fresh sandboxes.
-
Reconstructing Training Trajectories: After successfully completing tasks, the RSR framework reconstructs these solutions into training trajectories for supervised fine-tuning. This approach not only preserves the essence of successful solutions but also enhances the model's generalization capabilities through diverse experiences.
-
Experimental Results: In experiments, the RSR framework demonstrated superior performance across multiple benchmarks. For example, the pass@3 metric on Terminal-Bench 2 increased from 57.0% to 74.2%; on Terminal-Bench 4, from 1.5% to 9.1%; on the self-curated Terminal-Bench Hard, from 39.0% to 63.0%; and on the Software Terminal-Bench, from 3.0% to 6.0%. Additionally, the process reward on Long-Horizon Terminal-Bench rose from 0.21 to 0.29.
Technical Highlights
- Diverse Task-Solving Strategies: By employing task-solving strategies under different constraints, the RSR framework generates diverse training trajectories, enhancing the model's generalization capabilities.
- Reconstructing Training Trajectories: Successful solutions are reconstructed into training trajectories for supervised fine-tuning, ensuring the model learns from diverse experiences.
- Significant Performance Improvements: Experimental results show that the RSR framework significantly outperforms traditional methods, demonstrating its strong capabilities in complex tasks.
Industry Impact and Developer Recommendations
The introduction of the RSR framework provides AI researchers and developers with a new approach to enhancing model performance in complex tasks through diverse task-solving strategies. Developers can experiment with applying the RSR framework in their projects to improve the model's generalization and performance. Furthermore, the research findings from the RSR framework offer new insights into model training and performance enhancement, driving the advancement of AI technology.
Conclusion
Qwen's Recursive Self-Rewrite (RSR) framework demonstrates the potential of enhancing model capabilities through diverse task-solving experiences, paving the way for new directions in AI research.
— END —Source: Hugging Face Daily Papers (2026-10-02)
Tags: #Qwen #RSR #Large Language Model #Complex Task Solving #Supervised Fine-tuning
Community Comments