ZICQ
中 Log in / Sign up
ZICQ Info LLMs & Foundation Models #Nonobench #LLMs & Foundation Models #Reasoning #Open Source AI #AI Evaluation

Nonobench: An Open Benchmark for Evaluating 49 LLMs on Nonogram Puzzles

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Nonobench is an open-source benchmark tool designed to evaluate the reasoning capabilities of Large Language Models (LLMs) in solving nonogram (picross) puzzles. It assesses how well models can complete grids based on row and column clues without external tools, in both standard and hard modes, covering puzzles from 5x5 to 20x20. Results indicate a significant drop in performance as puzzle complexity increases, with GPT-6 Astra excelling in standard mode but struggling in hard mode. Nonobench pr


Overview

Nonobench is a newly released open-source tool designed to evaluate the reasoning capabilities of Large Language Models (LLMs) in solving nonogram (picross) puzzles. It assesses how well models can complete grids based on row and column clues without any external tools. Here are the key features and technical details of Nonobench:

Key Features

  1. Test Modes

    • Standard Mode: Includes 30 puzzles ranging from 5x5 to 15x15, sourced from Moyà-Alcover's Nonograms dataset (CC BY 4.0).
    • Hard Mode: Consists of 10 randomly generated 20x20 puzzles, with 5 that cannot be solved by line logic alone.
  2. Evaluation Method

    • Models are given only one attempt per puzzle and are not allowed to use external tools.
    • Results are presented as success rates, for example, GPT-6 Astra solved all 30 puzzles in standard mode.
  3. Result Analysis

    • The solve rate drops significantly as puzzle complexity increases, from 85% for 5x5 to 46% for 10x10, and 20% for 15x15.
    • In hard mode, Claude Opus 5.5 solved 8 puzzles, while 11 models solved none.
  4. Technical Details

    • Tests are run through OpenRouter and pinned to each lab's endpoint where possible.
    • To prevent models from guessing the picture content, random fills are used for puzzle grids.

Industry Impact

Nonobench provides AI researchers and developers with a platform to quantify and compare the reasoning abilities of LLMs, filling a gap in current LLM evaluation tools for complex logical tasks. Here are some key impacts:

  • Revealing Model Limitations: By testing models on puzzles of varying difficulty, Nonobench reveals the limitations of current LLMs in handling complex logical tasks.
  • Driving Model Improvement: The tool can help developers identify shortcomings in model reasoning capabilities, thereby driving improvements in model architecture and training methods.
  • Promoting Open Research: As an open-source tool, Nonobench encourages community participation and contribution, fostering open research and collaboration in the AI field.

Developer Recommendations

  • Use Nonobench for Model Evaluation: AI developers are encouraged to use Nonobench during model training and evaluation to quantify model reasoning capabilities.
  • Focus on Model Performance in Hard Mode: The results in hard mode can more accurately reflect the model's performance in complex logical tasks.
  • Participate in Community Discussion and Contribution: Developers can participate in Nonobench's community discussions, share test results and experiences, and contribute code to improve the tool.

Conclusion

Nonobench is an innovative open-source tool that provides a new method and perspective for evaluating the reasoning abilities of LLMs. By providing a standardized testing platform, Nonobench not only helps developers better understand model reasoning capabilities but also drives progress in the AI field in handling complex logical tasks.


Source: Reddit r/MachineLearning (2026-10-04)

— END —

Tags: #Nonobench #LLMs & Foundation Models #Reasoning #Open Source AI #AI Evaluation

Community Comments

Loading live comments and annotations…