Skip to main content
ZICQ

Overall LLM Rankings

ZICQ LLM rankings update daily across intelligence, coding, price, and context for selection and comparison.

Weighted composite of OpenRouter benchmarks Intelligence / Coding / Agentic (60% / 25% / 15%), covering all 338+ models.

Last updated: 29 min ago (2026-10-04 19:36)
Models tracked 466
Free & tunable 22
Vision 295
Tool calling 398
Reasoning 333
Max context Auto Router (Beta) 2M tokens
Avg price (input) — USD / 1M tokens
Cheapest model DeepSeek: DeepSeek V4.1 Flash $0.003 / 1M
TOP by Overall score Anthropic: Claude Opus 5.5 57.6
Clear
  1. #81
    SpaceXAI: Grok 4.7 xAI Doc & Vision Benchmarked

    Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and...

    Tools Vision Docs Reasoning Structured
    Intelligence 46.4
    Coding 0.0
    Agentic 0.0
    Context 500K Output ≤ 450K
    $2.00/1M $6.00/1M out cache $0.50/1M
    Recommended
  2. #82
    Anthropic: Claude Sonnet 4.5 Anthropic Doc & Vision Benchmarked

    Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

    Tools Vision Docs Reasoning Structured
    Intelligence 20.7
    Coding 52.1
    Agentic 15.8
    Context 1M Output ≤ 64K
    $3.00/1M $15.00/1M out cache $0.30/1M
    1M Context Pick
  3. #83
    Anthropic: Claude Sonnet 4.5 (batch) Anthropic Doc & Vision Benchmarked

    Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

    Tools Vision Docs Reasoning Structured
    Intelligence 20.7
    Coding 52.1
    Agentic 15.8
    Context 1M Output ≤ 64K
    $1.50/1M $7.50/1M out cache $0.15/1M
    1M Context Pick
  4. #84
    SpaceXAI: Grok 4.3 xAI Doc & Vision Benchmarked

    Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

    Tools Vision Docs Reasoning Structured
    Intelligence 24.9
    Coding 42.2
    Agentic 15.5
    Context 1M Output ≤ 900K
    $1.25/1M $2.50/1M out cache $0.20/1M
    1M Context Pick
  5. #85
    SpaceXAI: Grok 4.3 (batch) xAI Doc & Vision Benchmarked

    Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

    Tools Vision Docs Reasoning Structured
    Intelligence 24.9
    Coding 42.2
    Agentic 15.5
    Context 1M Output ≤ 900K
    $1.00/1M $2.00/1M out cache $0.16/1M
    1M Context Pick
  6. #86
    Google: Gemini 3.5 Flash Lite Google DeepMind Audio-Video Multimodal Benchmarked

    Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

    Tools Vision Audio Video Docs Reasoning Structured
    Intelligence 22.2
    Coding 49.3
    Agentic 14.3
    Context 1.0M Output ≤ 65K
    $0.30/1M $2.50/1M out cache $0.03/1M
    1M Context Pick
  7. #87
    Google: Gemini 3.5 Flash Lite (batch) Google DeepMind Audio-Video Multimodal Benchmarked

    Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

    Tools Vision Audio Video Docs Reasoning Structured
    Intelligence 22.2
    Coding 49.3
    Agentic 14.3
    Context 1.0M Output ≤ 65K
    $0.15/1M $1.25/1M out cache $0.01/1M
    1M Context Pick
  8. #88
    Xiaomi: MiMo-V2.6-Pro Xiaomi Audio-Video Multimodal Benchmarked

    MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding...

    Tools Vision Audio Video Reasoning Structured
    Intelligence 46.3
    Coding 0.0
    Agentic 0.0
    Context 1.1M Output ≤ 131K
    $0.43/1M $0.87/1M out cache $0.00/1M
    1M Context Pick
  9. #89
    inclusionAI: Ling 3.0 Flash Inclusionai General LLM Benchmarked

    *Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

    Tools Reasoning Structured
    Intelligence 20.1
    Coding 50.6
    Agentic 19.3
    Context 262K Output ≤ 32K
    $0.02/1M $0.06/1M out cache $0.00/1M
    High Value Pick
  10. #90
    Meituan: LongCat 2.0 Meituan General LLM Benchmarked

    LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...

    Tools Reasoning
    Intelligence 19.1
    Coding 45.3
    Agentic 14.0
    Context 1.0M Output ≤ 262K
    $0.30/1M $1.20/1M out cache $0.01/1M
    1M Context Pick
  11. #91
    Qwen: Qwen3.5 397B A17B Alibaba 通义 视觉多模态 Benchmarked

    The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

    Tools Vision Video Reasoning Structured
    Intelligence 18.4
    Coding 48.2
    Agentic 8.3
    Context 262K Output ≤ 235K
    $0.55/1M $3.50/1M out cache $0.22/1M
    Recommended
  12. #92
    Anthropic: Claude Opus 4.7 Anthropic Doc & Vision Benchmarked

    Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

    Tools Vision Docs Reasoning Structured
    Intelligence 0.0
    Coding 73.6
    Agentic 38.6
    Context 1M Output ≤ 128K
    $5.00/1M $25.00/1M out cache $0.50/1M
    1M Context Pick
  13. #93
    Anthropic: Claude Opus 4.7 (batch) Anthropic Doc & Vision Benchmarked

    Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

    Tools Vision Docs Reasoning Structured
    Intelligence 0.0
    Coding 73.6
    Agentic 38.6
    Context 1M Output ≤ 128K
    $2.50/1M $12.50/1M out cache $0.25/1M
    1M Context Pick
  14. #94
    DeepSeek: DeepSeek V4.1 Flash DeepSeek 视觉多模态 Benchmarked

    DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

    Tools Vision Reasoning Structured
    Intelligence 39.5
    Coding 0.0
    Agentic 0.0
    Context 1.0M Output ≤ 943K
    $0.00/1M $2.40/1M out cache $0.00/1M
    1M Context Pick
  15. #95
    DeepSeek: DeepSeek V4.1 Flash (batch) DeepSeek 视觉多模态 Benchmarked

    DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

    Tools Vision Reasoning Structured
    Intelligence 39.5
    Coding 0.0
    Agentic 0.0
    Context 1.0M Output ≤ 131K
    $0.11/1M $0.34/1M out cache $0.00/1M
    1M Context Pick
  16. #96
    Qwen: Qwen3.6 35B A3B Alibaba 通义 视觉多模态 Benchmarked

    Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

    Tools Vision Video Reasoning Structured
    Intelligence 18.2
    Coding 41.9
    Agentic 13.1
    Context 262K Output ≤ 235K
    $0.15/1M $1.00/1M out cache $0.05/1M
    High Value Pick
  17. #97
    OpenAI: GPT-6 Luna OpenAI Doc & Vision Benchmarked

    GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...

    Tools Vision Docs Reasoning Structured
    Intelligence 38.1
    Coding 0.0
    Agentic 0.0
    Context 1.1M Output ≤ 128K
    $0.10/1M $0.50/1M out cache $0.01/1M
    1M Context Pick
  18. #98
    OpenAI: GPT-6 Luna (batch) OpenAI Doc & Vision Benchmarked

    GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...

    Tools Vision Docs Reasoning Structured
    Intelligence 38.1
    Coding 0.0
    Agentic 0.0
    Context 1.1M Output ≤ 128K
    $0.05/1M $0.25/1M out cache $0.01/1M
    1M Context Pick
  19. #99
    Xiaomi: MiMo-V2.6-Flash Xiaomi Audio-Video Multimodal Benchmarked

    MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for...

    Tools Vision Audio Video Reasoning Structured
    Intelligence 37.9
    Coding 0.0
    Agentic 0.0
    Context 1.1M Output ≤ 131K
    $0.14/1M $0.28/1M out cache $0.00/1M
    1M Context Pick
  20. #100
    Anthropic: Claude Haiku 4.5 Anthropic Doc & Vision Benchmarked

    Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

    Tools Vision Docs Reasoning Structured
    Intelligence 16.9
    Coding 43.9
    Agentic 8.0
    Context 200K Output ≤ 64K
    $1.00/1M $5.00/1M out cache $0.10/1M
    Recommended