OpenRouter's own leaderboard is a usage board. It ranks models by the volume of tokens routed through OpenRouter over a window, so it tells you what developers are actually running in production right now. That is genuinely useful signal, but it is a popularity measure, not a quality one. A cheap, fast model wired into a high-volume app can outrank a far more capable model that people reserve for hard steps.
Quality is a separate axis. The scores you probably mean when you say "which LLM is best" come from independent evaluations, chiefly the Artificial Analysis indexes for intelligence, coding, and agentic ability. Those are measured on fixed test sets, not on how many people happen to be calling the model. This page keeps the two apart on purpose: usage tells you what is trusted, benchmarks tell you what is capable, and price tells you what you can afford to run at that capability.
The practical read: treat the raw usage rank as a shortlist of models people rely on, then decide with the benchmark board that matches your task and the price per million tokens you are willing to pay. The boards below do exactly that split.