Agents & Tool Calling
ZICQ LLM rankings update daily across intelligence, coding, price, and context for selection and comparison.
Ranked by OpenRouter benchmarks.agentic_index (agent / tool-calling index) plus the supported_parameters.tools flag — a reference for agent framework builders.
Last updated: 8 min ago (2026-10-03 15:30)-
#1
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
Tools Vision Docs Reasoning StructuredIntelligence 53.4Agentic 57.9Context 1M Output ≤ 128K$10.00/1M $50.00/1M out cache $0.25/1MTop Agentic Model -
#2
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
Tools Vision Docs Reasoning StructuredIntelligence 53.4Agentic 57.9Context 1M Output ≤ 128K$5.00/1M $25.00/1M out cache $0.12/1MAgent Framework Pick -
#3
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Tools Vision Docs Reasoning StructuredIntelligence 50.8Agentic 56.5Context 1M Output ≤ 128K$5.00/1M $25.00/1M out cache $0.50/1MAgent Framework Pick -
#4
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Tools Vision Docs Reasoning StructuredIntelligence 50.8Agentic 56.5Context 1M Output ≤ 128K$2.50/1M $12.50/1M out cache $0.25/1MSolid Tool Calling -
#5
Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...
Tools Vision Video Reasoning StructuredIntelligence 45.4Agentic 56.0Context 1M Output ≤ 131K$2.00/1M $6.00/1M out cache $0.25/1MSolid Tool Calling -
#6
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
Tools Reasoning StructuredIntelligence 44.8Agentic 53.1Context 1.0M Output ≤ 131K$1.40/1M $4.40/1M out cache $0.14/1MSolid Tool Calling -
#7
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
Tools Reasoning StructuredIntelligence 44.8Agentic 53.1Context 1.0M Output ≤ 131K$0.45/1M $2.00/1M out cache $0.10/1MSolid Tool Calling -
#8
Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by [Grok 4.7](/x-ai/grok-4.7).
Tools Vision Docs Reasoning StructuredIntelligence 44.3Agentic 53.0Context 500K Output ≤ 450K$2.00/1M $6.00/1M out cache $0.50/1MSolid Tool Calling -
#9
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
Tools Vision Docs Reasoning StructuredIntelligence 52.7Agentic 51.0Context 1.1M Output ≤ 128K$10.00/1M $50.00/1M out cache $1.00/1MSolid Tool Calling -
#10
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
Tools Vision Docs Reasoning StructuredIntelligence 52.7Agentic 51.0Context 1.1M Output ≤ 128K$5.00/1M $25.00/1M out cache $0.50/1MSolid Tool Calling -
#11
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Tools Vision Video Reasoning StructuredIntelligence 41.8Agentic 50.9Context 1.0M Output ≤ 943K$0.15/1M $0.50/1M out cache $0.03/1MSolid Tool Calling -
#12
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Tools Vision Video Reasoning StructuredIntelligence 41.8Agentic 50.9Context 1.0M Output ≤ 131K$0.06/1M $0.20/1M out cache $0.01/1MSolid Tool Calling -
#13
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Tools Vision Docs Reasoning StructuredIntelligence 49.6Agentic 50.7Context 1M Output ≤ 128K$10.00/1M $50.00/1M out cache $1.00/1MSolid Tool Calling -
#14
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Tools Vision Docs Reasoning StructuredIntelligence 49.6Agentic 50.7Context 1M Output ≤ 128K$5.00/1M $25.00/1M out cache $0.50/1MSolid Tool Calling -
#15
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
Tools Vision Docs Reasoning StructuredIntelligence 47.0Agentic 50.2Context 1.1M Output ≤ 128K$2.00/1M $10.00/1M out cache $0.20/1MSolid Tool Calling -
#16
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
Tools Vision Docs Reasoning StructuredIntelligence 47.0Agentic 50.2Context 1.1M Output ≤ 128K$1.00/1M $5.00/1M out cache $0.10/1MSolid Tool Calling -
#17
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...
Tools Reasoning StructuredIntelligence 39.9Agentic 50.1Context 1.0M Output ≤ 131K$2.00/1M $6.00/1M out cache $0.25/1MSolid Tool Calling -
#18
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Tools Vision Video Reasoning StructuredIntelligence 43.6Agentic 50.0Context 1.0M Output ≤ 943K$2.70/1M $13.50/1M out cache $0.27/1MSolid Tool Calling -
#19
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Tools Vision Video Reasoning StructuredIntelligence 43.6Agentic 50.0Context 1.0M Output ≤ 16K$2.28/1M $11.40/1M out cache $0.23/1MSolid Tool Calling -
#20
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
Tools Vision Video Reasoning StructuredIntelligence 33.7Agentic 45.8Context 1M Output ≤ 131K$0.42/1M $3.00/1M out cache $0.08/1MSolid Tool Calling