Enterprise DNA Enterprise DNA
Directories / Models / OpenRouter leaderboard

Models · Guide

OpenRouter leaderboard & LLM rankings, explained

The leaderboard everyone links to answers one question well and a different one badly. Here is what it actually measures, the current top models split by the job you are hiring them for, and how to read the numbers for real business tasks instead of chasing a single rank.

Updated

1342 models · OpenRouter + Artificial Analysis

What the OpenRouter leaderboard actually measures

OpenRouter's own leaderboard is a usage board. It ranks models by the volume of tokens routed through OpenRouter over a window, so it tells you what developers are actually running in production right now. That is genuinely useful signal, but it is a popularity measure, not a quality one. A cheap, fast model wired into a high-volume app can outrank a far more capable model that people reserve for hard steps.

Quality is a separate axis. The scores you probably mean when you say "which LLM is best" come from independent evaluations, chiefly the Artificial Analysis indexes for intelligence, coding, and agentic ability. Those are measured on fixed test sets, not on how many people happen to be calling the model. This page keeps the two apart on purpose: usage tells you what is trusted, benchmarks tell you what is capable, and price tells you what you can afford to run at that capability.

The practical read: treat the raw usage rank as a shortlist of models people rely on, then decide with the benchmark board that matches your task and the price per million tokens you are willing to pay. The boards below do exactly that split.

The operator's read

How to read the rankings for business tasks

There is no single best model, and the leaderboard is not trying to name one. The frontier now spans roughly a 50x range in output price, from premium reasoning models to cheap high-throughput ones, and the right pick changes with the job. The question that actually matters is per task: what is the cost of being wrong here, and what is the cheapest model that clears that bar.

For high-stakes, low-volume steps, such as an architecture decision, a contract review, or an eval judge, the price of a mistake dwarfs the token cost, so buy the top of the reasoning board and do not think twice. For high-volume, low-stakes work, such as classification, extraction, or first-draft generation, a mid or budget model that clears the quality bar wins on economics, and the cheapest-capable board is where you shop.

That is model routing, and it is the real skill the leaderboard is pointing at. Match each step of a workflow to the cheapest model that is good enough for that step, rather than paying frontier rates for everything or under-serving the steps that carry the risk.

Top models by use case

Live picks pulled from the model catalogue below. Each board sorts on the benchmark that matches the job, not on a blended score. Every model links to its full entry with pricing and access notes.

Best LLMs for coding

Ranked on the Artificial Analysis Coding Index, the benchmark that tracks real multi-file engineering rather than single-function puzzles. This is the list to read before you pick the model behind a coding agent.

Full coding board →
# Model Provider Coding Index $/M in $/M out Context
1 GPT-5.6 Sol OpenAI 77 $5.00 $30.00 1.1M
2 GPT-5.6 Terra OpenAI 77 $2.50 $15.00 1.1M
3 Claude Fable 5 Anthropic 77 $10.00 $50.00 1M
4 Kimi K3 Moonshotai 76 $3.00 $15.00 1.0M
5 GPT-5.5 OpenAI 75 $5.00 $30.00 1.1M
6 Claude Opus 4.8 Anthropic 74 $5.00 $25.00 1M

Best LLMs for agentic work

Ranked on the Agentic Index, which measures tool use, planning, and staying on task across long chains. Agentic scores and raw intelligence do not always agree, so this board is separate on purpose.

Full agentic board →
# Model Provider Agentic Index $/M in $/M out Context
1 GPT-5.6 Sol OpenAI 54 $5.00 $30.00 1.1M
2 Claude Fable 5 Anthropic 53 $10.00 $50.00 1M
3 Kimi K3 Moonshotai 50 $3.00 $15.00 1.0M
4 GPT-5.6 Terra OpenAI 47 $2.50 $15.00 1.1M
5 Claude Opus 4.8 Anthropic 47 $5.00 $25.00 1M
6 Claude Opus 4.8 (Fast) Anthropic 47 $10.00 $50.00 1M

Best LLMs for reasoning

Ranked on the Intelligence Index, the broad quality composite. Use these for the high-stakes steps: architecture calls, eval judging, hard analysis, anything where a wrong answer is expensive.

Full reasoning board →
# Model Provider Intelligence Index $/M in $/M out Context
1 Claude Fable 5 Anthropic 60 $10.00 $50.00 1M
2 GPT-5.6 Sol OpenAI 59 $5.00 $30.00 1.1M
3 Kimi K3 Moonshotai 57 $3.00 $15.00 1.0M
4 Claude Opus 4.8 Anthropic 56 $5.00 $25.00 1M
5 Claude Opus 4.8 (Fast) Anthropic 56 $10.00 $50.00 1M
6 GPT-5.6 Terra OpenAI 55 $2.50 $15.00 1.1M

Cheapest capable models

The lowest input price among models that still carry a real benchmark, so you are shopping for value, not just for zero. This is the board for high-volume steps where a good-enough answer at a fraction of the cost is the whole point.

Full cheapest board →
# Model Provider $/M in $/M out Intelligence Context
1 Mistral Small 3.2 24B Mistral $0.100 $0.300 944 131K
2 Llama 4 Scout Meta $0.100 $0.300 10 10M
3 Gemini 2.5 Flash Lite Google $0.100 $0.400 11 1.0M
4 GPT-4o-mini OpenAI $0.150 $0.600 11 128K
5 Llama 4 Maverick Meta $0.200 $0.800 14 1.0M
6 DeepSeek V3 DeepSeek $0.200 $0.800 1154 131K

Longest context windows

Sorted by how much you can put in one prompt. This is what matters for whole-repo reviews, long documents, and agents that carry a lot of working state between steps.

Browse all models →
# Model Provider Context $/M in Intelligence
1 Llama 4 Scout Meta 10M $0.100 10
2 Qwen3 Coder 480B A35B Qwen 1.0M $0.300 1180
3 Gemini 2.5 Flash Google 1.0M $0.300 1138
4 Gemini 2.5 Pro Google 1.0M $1.25 26
5 Llama 4 Maverick Meta 1.0M $0.200 14
6 Gemini 2.5 Flash Lite Google 1.0M $0.100 11
Read the columns

The metrics that matter

Intelligence Index
A broad quality composite from Artificial Analysis. The closest single number to general capability. Use it for the reasoning board and for high-stakes steps.
Coding Index
Coding-specific evaluation that rewards real multi-file engineering over single-function puzzles. The number to trust when the model sits behind a coding agent.
Agentic Index
Tool use, planning, and staying on task across long chains. High intelligence does not guarantee a high agentic score, which is why we rank them separately.
$/M in and $/M out
Price per million input and output tokens, synced from OpenRouter. Output is usually the pricier side and the one that dominates cost on generation-heavy jobs.
Context window
How much you can fit in one prompt. Decisive for whole-repo work, long documents, and agents that carry heavy working state between steps.

Common questions

What is the OpenRouter leaderboard?

It is a ranking of AI models by how many tokens developers route through OpenRouter over a period. It shows what is popular and trusted in production, which is a usage signal rather than a direct quality score.

Is the most-used model the best model?

Not necessarily. Usage rewards cheap, fast models wired into high-volume apps. For capability you want the benchmark boards, chiefly the Intelligence, Coding, and Agentic indexes, matched to your specific task.

Which LLM is best for coding right now?

The top of the Coding Index board above, refreshed from Artificial Analysis. It changes as new models ship, which is exactly why this page pulls it live from the catalogue instead of hard-coding a name.

How often is this updated?

Pricing and benchmarks sync regularly from OpenRouter and Artificial Analysis. The stamp at the top of this page shows the latest sync date reflected in the boards, currently Jul 20, 2026.