ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Turkish #AI Benchmark #Decision-Making #HakemBench #Open-Source

HakemBench Released: First Turkish Decision-Making Benchmark for AI Models

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:HakemBench, developed by HakkimLab, is the first benchmark for Turkish language decision-making tasks. It covers seven tracks, including fact-checking, education, legal routing, moderation, spam and phishing, and customer support, with 2,346 items and 4,275 questions. The benchmark is fully open-sourced under CC BY 4.0 and evaluates models based on decision quality, calibration, and selective automation using geometric mean scores. While the top-performing model achieves a composite score of 0.8


Release of the First Turkish Decision-Making Benchmark

HakemBench, developed by HakkimLab, is the first benchmark for Turkish language decision-making tasks. It aims to provide a standardized evaluation of AI models' decision-making capabilities in the Turkish language. Key features include:

  • Multi-domain Coverage: Encompasses seven main areas, including fact-checking, education, legal routing, moderation, spam and phishing, and customer support.
  • Rich Dataset: Contains 2,346 items and 4,275 questions, including multiple-choice, yes/no, and rating questions.
  • Open-Source License: Fully open-sourced under CC BY 4.0, facilitating easy access for researchers and developers.

Technical Highlights and Evaluation Metrics

The evaluation metrics of HakemBench include:

  1. Decision Quality (Macro F1): Measures the overall accuracy of the model in classification tasks.
  2. Calibration (Normalized Brier Score): Assesses the accuracy of the model's predicted probabilities.
  3. Selective Automation (Normalized Area Under the Generalized Risk-Coverage Curve): Evaluates the model's performance in automated decision-making.

These metrics are combined using a geometric mean, and confidence intervals are generated through 2,000 bootstrap samples. Additionally, HakemBench conducts probing tests for option order, paraphrase, English translation, and substituted names to assess the model's robustness.

Performance and Limitations

In the evaluation of 16 models, the top-performing model achieved a composite score of 0.888, while HakkimLab's own model ranked seventh with a score of 0.660. It is worth noting that some test results were flagged due to training data biases, indicating a high level of rigor and transparency in HakemBench's evaluation process.

Industry Impact and Developer Recommendations

The release of HakemBench provides a crucial tool for the evaluation of Turkish language AI models, particularly in the context of multilingual AI applications. Here are some recommendations:

  • Model Optimization: Developers can use HakemBench to optimize models, especially for multi-domain decision-making tasks.
  • Data Quality Control: Pay attention to the quality of training data to avoid model performance issues due to data biases.
  • Cross-Language Applications: Explore the extension of HakemBench to other languages, promoting the development of multilingual AI technology.

Conclusion

The release of HakemBench marks a significant advancement in the evaluation of Turkish language AI models, providing valuable resources for researchers and developers. Despite some limitations, its open-source nature and multi-domain coverage make it an important tool for the AI community.


Source: ArXiv NLP/LLM (cs.CL) (2026-10-05)

— END —

Tags: #Turkish #AI Benchmark #Decision-Making #HakemBench #Open-Source

Community Comments

Loading live comments and annotations…