Hugging Face Introduces OC-SFT Framework to Enhance Decision Consistency in Large Language Models
By Mr.Xu
Published:
Summary:Hugging Face's research team introduces the Order-Consistency SFT (OC-SFT) framework to address the order dependence problem in Large Language Models (LLMs) when scoring multiple candidates. OC-SFT penalizes discrepancies in candidate scores across different orderings in the weights, significantly improving decision consistency while maintaining ranking quality. This method outperforms existing approaches in various benchmarks, demonstrating its potential in tasks like passage reranking, respons
Background and Challenge
In tasks like passage reranking, response scoring, and multi-document question answering, Large Language Models (LLMs) often need to score multiple candidate documents or responses simultaneously. While these scorers are selected based on ranking quality, their scores determine the final decision, such as what content to retain, the reader's answer, or which candidate pairs enter preference training. However, reordering the candidates changes their scores, leading to inconsistent decisions for the same query. This order dependence issue makes it difficult for LLMs to ensure the stability and reliability of decisions in practical applications.
Technical Breakthrough
Hugging Face's Order-Consistency SFT (OC-SFT) framework addresses this problem by penalizing discrepancies in candidate scores across different orderings in the weights. Specifically, OC-SFT introduces a new loss function during Supervised Fine-Tuning (SFT) that encourages the model to maintain consistent scores for the same candidate across different orderings. This method not only preserves the original ranking quality but also excels in multiple decision stability metrics.
Experimental Results
The experimental results show that OC-SFT performs excellently in the following aspects:
- Ranking Quality: OC-SFT maintains ranking quality comparable to existing methods in multiple benchmarks.
- Decision Stability: OC-SFT significantly improves decision stability across all three tasks. For example, in the passage reranking task, the overlap of the retained set by OC-SFT is over 10% higher than traditional methods.
- Model Robustness: OC-SFT demonstrates higher stability on 12 base models compared to the order-averaged distillation method.
Industry Impact and Developer Recommendations
The introduction of OC-SFT provides a new solution to the decision stability problem of LLMs in practical applications. Here are some recommendations:
- Application Scenario Expansion: Developers can apply OC-SFT in tasks such as text reranking, response scoring, and multi-document question answering to enhance the reliability of model decisions.
- Model Optimization: The loss function design of OC-SFT can inspire model optimization in other fields, such as the order dependence problem in multimodal models.
- Further Research: Future research can explore the combination of OC-SFT with other technologies, such as reinforcement learning and prompt engineering, to further improve LLM performance.
Conclusion
The OC-SFT framework significantly enhances the decision stability and consistency of LLMs by solving the order dependence problem, providing a new technical pathway for building more reliable intelligent systems.
— END —Source: Hugging Face Daily Papers (2026-09-26)
Tags: #Hugging Face #Large Language Model #Supervised Fine-Tuning #Decision Stability
Community Comments