ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #ArXiv #Large Language Models #Multi-Armed Bandit #Automated Scoring #Cost Optimization

ArXiv Proposes Bandit-Driven Framework to Optimize LLM Essay Scoring Efficiency

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:The ArXiv team proposes a novel multi-armed bandit (MAB) framework to optimize the prompt selection strategy for automated essay scoring (AES) using large language models (LLMs). This approach treats different prompt types as arms in an MAB controller, enabling adaptive selection of the optimal prompting strategy during inference. Experiments on IELTS Writing Task 2 essays demonstrate that the MAB framework achieves comparable scoring accuracy to exhaustive grid search while reducing LLM calls b


Key Breakthroughs

The ArXiv team has introduced a novel multi-armed bandit (MAB) framework for optimizing the prompt selection strategy in automated essay scoring (AES) using large language models (LLMs). Here are the key points of the research:

  1. Application of MAB Framework:

    • Different prompt types are treated as arms in an MAB controller, allowing the system to adaptively select the optimal prompt strategy during inference.
    • This approach dynamically adjusts the prompt selection based on real-time feedback during the reasoning process.
  2. Experimental Results:

    • The method achieves comparable scoring accuracy to exhaustive grid search in IELTS Writing Task 2 essays.
    • It significantly reduces LLM calls, lowering computational costs by 78.4%.
  3. Cost-Reliability Learning Curves:

    • The study produces the first cost-reliability learning curves for essay scoring, offering practical insights for educational technology platforms to balance operational costs and assessment validity.
  4. Advantages of Multi-Step Scoring Strategy:

    • The research finds that the multi-step scoring strategy with examples outperforms single-step and no-example strategies in terms of accuracy.

Technical Highlights

  • Adaptive Prompt Selection: The MAB framework enables the model to dynamically adjust the prompt selection strategy based on real-time feedback during the reasoning process, overcoming the limitations of traditional fixed prompt selection strategies.
  • Efficient Resource Utilization: The method significantly reduces LLM calls, lowering computational costs while maintaining high scoring accuracy.
  • Data-Driven Decision Support: The generated cost-reliability learning curves provide data-driven decision support for educational technology platforms, helping them find the optimal balance between cost and assessment accuracy.

Industry Impact

  • Optimization of Educational Technology Platforms: The method offers an efficient and cost-effective solution for automated scoring in educational technology platforms, helping to improve scoring efficiency and reduce operational costs.
  • Expansion of LLM Application Scenarios: By optimizing the prompt selection strategy, the study demonstrates the potential of LLMs in complex tasks, providing new ideas for their application in more fields.

Developer Recommendations

  • Implement MAB Framework: Developers should consider implementing the MAB framework in LLM applications to optimize the prompt selection strategy and improve system efficiency.
  • Focus on Cost-Reliability Balance: When designing automated scoring systems, attention should be paid to the balance between cost and reliability, and strategies should be adjusted according to specific application scenarios.
  • Continuously Optimize the Model: Combine multi-step scoring strategies and example data to continuously optimize the model and improve scoring accuracy.

Source: ArXiv cs.AI (2026-08-24)

— END —

Tags: #ArXiv #Large Language Models #Multi-Armed Bandit #Automated Scoring #Cost Optimization

Community Comments

Loading live comments and annotations…