OpenRouter Rankings July 2026
Who's Actually Winning the AI Model Race?

Paid token traffic · China ~46% share · volume ≠ quality · agent app layer · August outlook

OpenRouter July 2026 AI model rankings vendor share chart
If your model picks still come from last quarter's headlines, the July chart will feel like a different industry. OpenRouter is the largest neutral routing layer for LLM APIs — its leaderboard measures where developers actually send paid traffic, not benchmark scores. This guide, grounded in data through July 25, 2026, covers: ① the Top 12 models and vendor share reversal; ② the barbell split between volume leaders and quality spend; ③ Hermes Agent, Kilo Code, and the hidden roleplay traffic stack; ④ pricing comparison and scenario matrix; ⑤ five August trend calls plus a six-step tiered routing runbook.
01

How to read the July OpenRouter leaderboard: Xiaomi on top and China at ~46%

OpenRouter ranks models by token volume — a live tally of which APIs developers pay to hit at scale. As of July 25, the daily leaders were Xiaomi Mimo V2.5 (1.4T/day), DeepSeek V4 Flash (943.9B/day), and Tencent Hy3 (590B/day). Seven of the top ten models come from Chinese labs; only Nemotron 3 Ultra (NVIDIA), Claude, and Gemini entries hold US-side positions in the upper tier.

RankModelVendorDaily tokens30-day total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot AI157.6B1.6T (new entry)
10Ling 3.0 FlashAnt InclusionAI128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

Zoom out to vendor share: Chinese labs collectively hold about 46% of token volume, up from under 2% a year ago. The US big three (OpenAI, Anthropic, Google) slid from roughly 70% at mid-2025 to 30%–36%. The driver is arithmetic, not narrative — DeepSeek V4 Flash input runs $0.05–0.14/M versus GPT-5.5 around $5/M, a 35× gap on comparable list pricing.

01

DeepSeek is the steady anchor: single-vendor share sits around 16%–18%, often #1 by company even when the daily model crown rotates.

02

Xiaomi swings hardest: Mimo V2.5 spikes can move vendor share between 8% and 18% within weeks.

03

Kimi K3 is the fastest climber: new to the board and already Top 9, backed by a record 1.4TB open-weight release.

04

Daily noise is real: Claude Opus 4.8 still ranked #10 on 7/24 but fell off the Top 12 on 7/25 when Ling 3.0 Flash moved up.

05

Track persistence, not snapshots: who tops Tuesday matters less than who stays in the front row for a month.

02

Can leaderboard rank equal quality? The volume–capability barbell

A necessary caveat: OpenRouter sorts by tokens, not intelligence. Cheap, fast models wired behind high-volume apps can dominate the chart even when they stumble on hard reasoning — and July proves the market is splitting along that fault line.

DimensionBy token volumeBy spend on hard tasks
Top tierChinese open models at aggressive price pointsClaude Sonnet 4.6 / Opus 4.7 tied at 13.5% each
Typical workloadChat, creative writing, roleplay, light codingComplex reasoning, enterprise agent planning, high-consistency classification
Price logicThroughput, fault tolerance, 35× spread vs GPT-5.5Closed frontier models keep $5/$25 per-M pricing power
July exampleMimo V2.5 at 1.4T/dayClaude Opus 5 at 43.3% on FrontierBench v0.1

The market is self-segmenting: inexpensive Chinese open models absorb bulk low-bar tasks, while closed frontier SKUs still own hard-work pricing and security reputation.

Spend mix by task type shows general dialogue at 35.7%, agent workflows 30.4%, code 26.5%, and data processing 7.5%. Narrow the lens to classification and complex reasoning and GPT-5.5 holds 11.6% — third place — while volume-board cheap open models barely register. Anthropic's July 24 Claude Opus 5 launch scored 43.3% on FrontierBench v0.1 (GPT-5.6 Sol at 37.5%) while keeping Opus-tier $5/$25 per-M list pricing — a blunt "expensive, but worth it on the hardest slice" message.

ModelInput/MOutput/MPositioning
DeepSeek V4 Flash$0.05–0.14$0.24–0.28Cost king, agentic coding default
GLM 5.2$0.45$3.31Near-Opus planning among open weights
MiniMax M3$0.10$1.21Long-context multimodal budget pick
Kimi K3~$3~$151.4TB open weights, closed-tier capability
Claude Opus 5$5 (fast tier $10)$25 (fast tier $50)Closed frontier, top July benchmark score
03

OpenRouter app rankings: coding agents lead, roleplay traffic stays invisible

Model charts show which "brains" get called; the app layer at openrouter.ai/apps shows what those brains actually do in production.

RankAppTypeShare (approx.)
1Hermes AgentPersonal / CLI agent~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude CodeCoding agent~6%
5DescriptContent production~4.5%
6–10pi, Lemonade, ISEKAI ZERO, Janitor AI, ClineAgent / roleplay1.7%–3.3%

Key read: Cline → Roo Code → Kilo Code is one coding-agent lineage forked three times — the "grandchild" Kilo Code has overtaken its ancestors. OpenRouter and a16z's State of AI report show creative roleplay driving more than half of open-model usage, a workload category enterprise AI coverage barely mentions.

ScenarioModel pickWhy
Daily coding / agentic devDeepSeek V4 FlashBest price-performance; #2 on OpenRouter volume
Open planning qualityGLM 5.2Closest open-weight match to Opus-style planning
Hard reasoning / classificationClaude Opus 5 / Sonnet 5Lead spend share on difficult tasks
Long-context open stackKimi K31M context, 1.4TB weights (pre-quant, cloud inference today)
Multimodal on a budgetMiniMax M3Strong image-input economics
US full open stackNemotron 3 UltraNVIDIA ecosystem, free tier available
04

How to pick an AI model API: six-step tiered routing runbook

01

Grade your workloads: split traffic into bulk tier (chat, summaries, simple completions) and hard tier (complex reasoning, high-risk agent decisions). Never route everything through one model SKU.

02

Default bulk to DeepSeek V4 Flash: $0.05–0.14/M input, agentic coding workhorse; use GLM 5.2 when open-weight planning needs a step up.

03

Reserve Claude Opus 5 for hard tier: leads classification and complex-reasoning spend; escalate only after cheaper models fail twice — true hybrid routing.

04

Wire OpenRouter as one endpoint: single key across hundreds of models for R&D and A/B tests; budget ~180–250ms added latency from many regions.

05

Add security to the scorecard: this week's OpenAI sandbox-escape headlines matter for enterprise agents — tighten permissions and review vendor safety track records.

06

Host routing 24/7: run agent gateways on always-on cloud Mac Mini nodes so long jobs survive laptop sleep; evaluate compliant production relays for your market.

OpenRouter hybrid routing example
curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -d '{
    "models": [
      "deepseek/deepseek-v4-flash",
      "anthropic/claude-opus-5"
    ],
    "messages": [{"role": "user", "content": "Plan a multi-step refactor..."}]
  }'
05

August 2026 OpenRouter outlook and three hard data points

Extrapolating July momentum, here is what August likely brings:

1

Chinese open-source share climbs toward 50% unless US vendors cut list prices materially — no sign of that yet.

2

Monthly model champions keep rotating: Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are not slowing price-or-iteration cycles.

3

Anthropic may ship a cheaper tier: Opus 5 is the fourth flagship in two months — a multi-price-band strategy, not a single-hero launch.

4

Kimi K3 community quants land in 2–4 weeks: 1.4TB weights need compression before most teams self-host; inference clouds benefit today.

5

Security enters procurement checklists: US congressional AI kill-switch legislation and White House frontier-model pre-release review frameworks may finalize before August ends.

A

Chinese vendor token share: about 46% (under 2% a year ago) — one of the steepest twelve-month share migrations in AI.

B

Hermes Agent app share: about 45% — largest single application on OpenRouter.

C

DeepSeek vs GPT-5.5 input pricing: roughly 35× gap ($0.05–0.14/M vs ~$5/M), fueling bulk-task migration.

Note: rankings shift daily — confirm the latest Top 10 at openrouter.ai/rankings before you publish internal picks.

The July line to remember: capability and popularity are diverging. Time spent arguing "who is #1" beats time spent building evaluation harnesses and tiered routing — and routing layers need hosts that stay online.

Laptop gateways lose fights with sleep mode, RAM ceilings, and Wi-Fi jitter; shared VPS CPUs struggle with Xcode and Metal inference stability. Teams running Hermes Agent, OpenClaw, or multi-model CI around the clock benefit from MESHLAUNCH bare-metal Mac Mini cloud rental — dedicated Apple Silicon, flexible daily/weekly/monthly terms, and production-grade uptime. See the pricing page and help center for regions and setup.

FAQ

As of July 25, by daily token volume, Xiaomi Mimo V2.5 led at 1.4 trillion tokens per day, followed by DeepSeek V4 Flash (943.9B) and Tencent Hy3 (590B). Kimi K3 entered at #9 — the fastest new climber on the board.

No. The board sorts by token volume, which favors cheap high-throughput workloads. On hard reasoning, Claude Sonnet 4.6 and Opus 4.7 each hold 13.5% spend share. For stable agent hosting, see the pricing page.

Start with DeepSeek V4 Flash and GLM 5.2 for coding, escalate stuck steps to Claude Opus 5. OpenRouter works well as an R&D sandbox; production needs latency and compliance review for your region.

Deploy OpenRouter or LiteLLM routing on an always-on cloud Mac. Region and networking guidance lives in the help center; match daily or monthly rental to project length.