Enterprise DNA Enterprise DNA
Directories / Models / OpenRouter leaderboard

Models · Guide

OpenRouter leaderboard & LLM rankings, explained

The leaderboard everyone links to answers one question well and a different one badly. Here is what it actually measures, the current top models split by the job you are hiring them for, and how to read the numbers for real business tasks instead of chasing a single rank.

Updated

1465 models · OpenRouter + Artificial Analysis

What the OpenRouter leaderboard actually measures

OpenRouter's own leaderboard is a usage board. It ranks models by the volume of tokens routed through OpenRouter over a window, so it tells you what developers are actually running in production right now. That is genuinely useful signal, but it is a popularity measure, not a quality one. A cheap, fast model wired into a high-volume app can outrank a far more capable model that people reserve for hard steps.

Quality is a separate axis. The scores you probably mean when you say "which LLM is best" come from independent evaluations, chiefly the Artificial Analysis indexes for intelligence, coding, and agentic ability. Those are measured on fixed test sets, not on how many people happen to be calling the model. This page keeps the two apart on purpose: usage tells you what is trusted, benchmarks tell you what is capable, and price tells you what you can afford to run at that capability.

The practical read: treat the raw usage rank as a shortlist of models people rely on, then decide with the benchmark board that matches your task and the price per million tokens you are willing to pay. The boards below do exactly that split.

The operator's read

How to read the rankings for business tasks

There is no single best model, and the leaderboard is not trying to name one. The frontier now spans roughly a 50x range in output price, from premium reasoning models to cheap high-throughput ones, and the right pick changes with the job. The question that actually matters is per task: what is the cost of being wrong here, and what is the cheapest model that clears that bar.

For high-stakes, low-volume steps, such as an architecture decision, a contract review, or an eval judge, the price of a mistake dwarfs the token cost, so buy the top of the reasoning board and do not think twice. For high-volume, low-stakes work, such as classification, extraction, or first-draft generation, a mid or budget model that clears the quality bar wins on economics, and the cheapest-capable board is where you shop.

That is model routing, and it is the real skill the leaderboard is pointing at. Match each step of a workflow to the cheapest model that is good enough for that step, rather than paying frontier rates for everything or under-serving the steps that carry the risk.

Top models by use case

Live picks pulled from the model catalogue below. Each board sorts on the benchmark that matches the job, not on a blended score. Every model links to its full entry with pricing and access notes.

Best LLMs for coding

Ranked on the Artificial Analysis Coding Index, the benchmark that tracks real multi-file engineering rather than single-function puzzles. This is the list to read before you pick the model behind a coding agent.

Full coding board →
# Model Provider Coding Index $/M in $/M out Context
1 Claude Fable 5.1 (batch) Anthropic 82 $5.00 $25.00 1M
2 Claude Fable 5.1 Anthropic 82 $10.00 $50.00 1M
3 Claude Opus 5 (batch) Anthropic 78 $2.50 $12.50 1M
4 Claude Opus 5 Anthropic 78 $5.00 $25.00 1M
5 GPT-5.6 Sol (batch) OpenAI 77 $1.00 $5.00 1.1M
6 GPT-5.6 Sol OpenAI 77 $2.00 $10.00 1.1M

Best LLMs for agentic work

Ranked on the Agentic Index, which measures tool use, planning, and staying on task across long chains. Agentic scores and raw intelligence do not always agree, so this board is separate on purpose.

Full agentic board →
# Model Provider Agentic Index $/M in $/M out Context
1 Claude Fable 5.1 (batch) Anthropic 58 $5.00 $25.00 1M
2 Claude Fable 5.1 Anthropic 58 $10.00 $50.00 1M
3 Claude Opus 5 (batch) Anthropic 56 $2.50 $12.50 1M
4 Claude Opus 5 Anthropic 56 $5.00 $25.00 1M
5 Grok 4.6 xAI 54 $2.00 $6.00 500K
6 GLM 5.3 Z Ai 54 $1.40 $4.40 1.3M

Best LLMs for reasoning

Ranked on the Intelligence Index, the broad quality composite. Use these for the high-stakes steps: architecture calls, eval judging, hard analysis, anything where a wrong answer is expensive.

Full reasoning board →
# Model Provider Intelligence Index $/M in $/M out Context
1 Claude Fable 5.1 (batch) Anthropic 57 $5.00 $25.00 1M
2 Claude Fable 5.1 Anthropic 57 $10.00 $50.00 1M
3 Claude Opus 4.8 (Fast) Anthropic 56 $10.00 $50.00 1M
4 GPT-6 Astra (batch) OpenAI 55 $5.00 $25.00 1.1M
5 GPT-6 Astra OpenAI 55 $10.00 $50.00 1.1M
6 Claude Opus 5 (batch) Anthropic 54 $2.50 $12.50 1M

Cheapest capable models

The lowest input price among models that still carry a real benchmark, so you are shopping for value, not just for zero. This is the board for high-volume steps where a good-enough answer at a fraction of the cost is the whole point.

Full cheapest board →
# Model Provider $/M in $/M out Intelligence Context
1 Mistral Small 3.2 24B Mistral $0.075 $0.200 924 131K
2 Llama 4 Scout Meta $0.100 $0.300 8 1.3M
3 Gemini 2.5 Flash Lite Google $0.100 $0.400 11 1.0M
4 GPT-4o-mini OpenAI $0.150 $0.600 11 128K
5 Llama 4 Maverick Meta $0.200 $0.696 16 1.0M
6 DeepSeek V3.2 DeepSeek $0.269 $0.400 44 164K

Longest context windows

Sorted by how much you can put in one prompt. This is what matters for whole-repo reviews, long documents, and agents that carry a lot of working state between steps.

Browse all models →
# Model Provider Context $/M in Intelligence
1 Llama 4 Scout Meta 1.3M $0.100 8
2 Gemini 2.5 Flash Google 1.0M $0.300 1108
3 Gemini 2.5 Pro Google 1.0M $1.25 33
4 Llama 4 Maverick Meta 1.0M $0.200 16
5 Gemini 2.5 Flash Lite Google 1.0M $0.100 11
6 Claude Opus 4.6 Anthropic 1M $5.00 1226
Read the columns

The metrics that matter

Intelligence Index
A broad quality composite from Artificial Analysis. The closest single number to general capability. Use it for the reasoning board and for high-stakes steps.
Coding Index
Coding-specific evaluation that rewards real multi-file engineering over single-function puzzles. The number to trust when the model sits behind a coding agent.
Agentic Index
Tool use, planning, and staying on task across long chains. High intelligence does not guarantee a high agentic score, which is why we rank them separately.
$/M in and $/M out
Price per million input and output tokens, synced from OpenRouter. Output is usually the pricier side and the one that dominates cost on generation-heavy jobs.
Context window
How much you can fit in one prompt. Decisive for whole-repo work, long documents, and agents that carry heavy working state between steps.

Common questions

What is the OpenRouter leaderboard?

It is a ranking of AI models by how many tokens developers route through OpenRouter over a period. It shows what is popular and trusted in production, which is a usage signal rather than a direct quality score.

Is the most-used model the best model?

Not necessarily. Usage rewards cheap, fast models wired into high-volume apps. For capability you want the benchmark boards, chiefly the Intelligence, Coding, and Agentic indexes, matched to your specific task.

Which LLM is best for coding right now?

The top of the Coding Index board above, refreshed from Artificial Analysis. It changes as new models ship, which is exactly why this page pulls it live from the catalogue instead of hard-coding a name.

How often is this updated?

Pricing and benchmarks sync regularly from OpenRouter and Artificial Analysis. The stamp at the top of this page shows the latest sync date reflected in the boards, currently Sep 6, 2026.