How to read the July OpenRouter leaderboard: Xiaomi on top and China at ~46%
OpenRouter ranks models by token volume — a live tally of which APIs developers pay to hit at scale. As of July 25, the daily leaders were Xiaomi Mimo V2.5 (1.4T/day), DeepSeek V4 Flash (943.9B/day), and Tencent Hy3 (590B/day). Seven of the top ten models come from Chinese labs; only Nemotron 3 Ultra (NVIDIA), Claude, and Gemini entries hold US-side positions in the upper tier.
| Rank | Model | Vendor | Daily tokens | 30-day total |
|---|---|---|---|---|
| 1 | Mimo V2.5 | Xiaomi | 1.4T | 31.2T |
| 2 | DeepSeek V4 Flash | DeepSeek | 943.9B | 23.6T |
| 3 | Hy3 | Tencent | 590B | 23.4T |
| 4 | Nemotron 3 Ultra (free) | NVIDIA | 428.6B | 9T |
| 5 | DeepSeek V4 Pro | DeepSeek | 413.7B | 11.6T |
| 6 | GLM 5.2 | Z.ai | 316.7B | 13.3T |
| 7 | MiniMax M3 | MiniMax | 262.5B | 15.1T |
| 8 | Step 3.7 Flash | StepFun | 204.8B | 5.9T |
| 9 | Kimi K3 | Moonshot AI | 157.6B | 1.6T (new entry) |
| 10 | Ling 3.0 Flash | Ant InclusionAI | 128.3B | 417.3B |
| 11 | Gemini 3 Flash Preview | 106.3B | 4T | |
| 12 | Claude Sonnet 5 | Anthropic | 99.5B | 3.6T |
Zoom out to vendor share: Chinese labs collectively hold about 46% of token volume, up from under 2% a year ago. The US big three (OpenAI, Anthropic, Google) slid from roughly 70% at mid-2025 to 30%–36%. The driver is arithmetic, not narrative — DeepSeek V4 Flash input runs $0.05–0.14/M versus GPT-5.5 around $5/M, a 35× gap on comparable list pricing.
DeepSeek is the steady anchor: single-vendor share sits around 16%–18%, often #1 by company even when the daily model crown rotates.
Xiaomi swings hardest: Mimo V2.5 spikes can move vendor share between 8% and 18% within weeks.
Kimi K3 is the fastest climber: new to the board and already Top 9, backed by a record 1.4TB open-weight release.
Daily noise is real: Claude Opus 4.8 still ranked #10 on 7/24 but fell off the Top 12 on 7/25 when Ling 3.0 Flash moved up.
Track persistence, not snapshots: who tops Tuesday matters less than who stays in the front row for a month.
Can leaderboard rank equal quality? The volume–capability barbell
A necessary caveat: OpenRouter sorts by tokens, not intelligence. Cheap, fast models wired behind high-volume apps can dominate the chart even when they stumble on hard reasoning — and July proves the market is splitting along that fault line.
| Dimension | By token volume | By spend on hard tasks |
|---|---|---|
| Top tier | Chinese open models at aggressive price points | Claude Sonnet 4.6 / Opus 4.7 tied at 13.5% each |
| Typical workload | Chat, creative writing, roleplay, light coding | Complex reasoning, enterprise agent planning, high-consistency classification |
| Price logic | Throughput, fault tolerance, 35× spread vs GPT-5.5 | Closed frontier models keep $5/$25 per-M pricing power |
| July example | Mimo V2.5 at 1.4T/day | Claude Opus 5 at 43.3% on FrontierBench v0.1 |
The market is self-segmenting: inexpensive Chinese open models absorb bulk low-bar tasks, while closed frontier SKUs still own hard-work pricing and security reputation.
Spend mix by task type shows general dialogue at 35.7%, agent workflows 30.4%, code 26.5%, and data processing 7.5%. Narrow the lens to classification and complex reasoning and GPT-5.5 holds 11.6% — third place — while volume-board cheap open models barely register. Anthropic's July 24 Claude Opus 5 launch scored 43.3% on FrontierBench v0.1 (GPT-5.6 Sol at 37.5%) while keeping Opus-tier $5/$25 per-M list pricing — a blunt "expensive, but worth it on the hardest slice" message.
| Model | Input/M | Output/M | Positioning |
|---|---|---|---|
| DeepSeek V4 Flash | $0.05–0.14 | $0.24–0.28 | Cost king, agentic coding default |
| GLM 5.2 | $0.45 | $3.31 | Near-Opus planning among open weights |
| MiniMax M3 | $0.10 | $1.21 | Long-context multimodal budget pick |
| Kimi K3 | ~$3 | ~$15 | 1.4TB open weights, closed-tier capability |
| Claude Opus 5 | $5 (fast tier $10) | $25 (fast tier $50) | Closed frontier, top July benchmark score |
OpenRouter app rankings: coding agents lead, roleplay traffic stays invisible
Model charts show which "brains" get called; the app layer at openrouter.ai/apps shows what those brains actually do in production.
| Rank | App | Type | Share (approx.) |
|---|---|---|---|
| 1 | Hermes Agent | Personal / CLI agent | ~45% |
| 2 | Kilo Code | Coding agent | ~13% |
| 3 | OpenClaw | General agent | ~9% |
| 4 | Claude Code | Coding agent | ~6% |
| 5 | Descript | Content production | ~4.5% |
| 6–10 | pi, Lemonade, ISEKAI ZERO, Janitor AI, Cline | Agent / roleplay | 1.7%–3.3% |
Key read: Cline → Roo Code → Kilo Code is one coding-agent lineage forked three times — the "grandchild" Kilo Code has overtaken its ancestors. OpenRouter and a16z's State of AI report show creative roleplay driving more than half of open-model usage, a workload category enterprise AI coverage barely mentions.
| Scenario | Model pick | Why |
|---|---|---|
| Daily coding / agentic dev | DeepSeek V4 Flash | Best price-performance; #2 on OpenRouter volume |
| Open planning quality | GLM 5.2 | Closest open-weight match to Opus-style planning |
| Hard reasoning / classification | Claude Opus 5 / Sonnet 5 | Lead spend share on difficult tasks |
| Long-context open stack | Kimi K3 | 1M context, 1.4TB weights (pre-quant, cloud inference today) |
| Multimodal on a budget | MiniMax M3 | Strong image-input economics |
| US full open stack | Nemotron 3 Ultra | NVIDIA ecosystem, free tier available |
How to pick an AI model API: six-step tiered routing runbook
Grade your workloads: split traffic into bulk tier (chat, summaries, simple completions) and hard tier (complex reasoning, high-risk agent decisions). Never route everything through one model SKU.
Default bulk to DeepSeek V4 Flash: $0.05–0.14/M input, agentic coding workhorse; use GLM 5.2 when open-weight planning needs a step up.
Reserve Claude Opus 5 for hard tier: leads classification and complex-reasoning spend; escalate only after cheaper models fail twice — true hybrid routing.
Wire OpenRouter as one endpoint: single key across hundreds of models for R&D and A/B tests; budget ~180–250ms added latency from many regions.
Add security to the scorecard: this week's OpenAI sandbox-escape headlines matter for enterprise agents — tighten permissions and review vendor safety track records.
Host routing 24/7: run agent gateways on always-on cloud Mac Mini nodes so long jobs survive laptop sleep; evaluate compliant production relays for your market.
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-d '{
"models": [
"deepseek/deepseek-v4-flash",
"anthropic/claude-opus-5"
],
"messages": [{"role": "user", "content": "Plan a multi-step refactor..."}]
}'
August 2026 OpenRouter outlook and three hard data points
Extrapolating July momentum, here is what August likely brings:
Chinese open-source share climbs toward 50% unless US vendors cut list prices materially — no sign of that yet.
Monthly model champions keep rotating: Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are not slowing price-or-iteration cycles.
Anthropic may ship a cheaper tier: Opus 5 is the fourth flagship in two months — a multi-price-band strategy, not a single-hero launch.
Kimi K3 community quants land in 2–4 weeks: 1.4TB weights need compression before most teams self-host; inference clouds benefit today.
Security enters procurement checklists: US congressional AI kill-switch legislation and White House frontier-model pre-release review frameworks may finalize before August ends.
Chinese vendor token share: about 46% (under 2% a year ago) — one of the steepest twelve-month share migrations in AI.
Hermes Agent app share: about 45% — largest single application on OpenRouter.
DeepSeek vs GPT-5.5 input pricing: roughly 35× gap ($0.05–0.14/M vs ~$5/M), fueling bulk-task migration.
Note: rankings shift daily — confirm the latest Top 10 at openrouter.ai/rankings before you publish internal picks.
The July line to remember: capability and popularity are diverging. Time spent arguing "who is #1" beats time spent building evaluation harnesses and tiered routing — and routing layers need hosts that stay online.
Laptop gateways lose fights with sleep mode, RAM ceilings, and Wi-Fi jitter; shared VPS CPUs struggle with Xcode and Metal inference stability. Teams running Hermes Agent, OpenClaw, or multi-model CI around the clock benefit from MESHLAUNCH bare-metal Mac Mini cloud rental — dedicated Apple Silicon, flexible daily/weekly/monthly terms, and production-grade uptime. See the pricing page and help center for regions and setup.
As of July 25, by daily token volume, Xiaomi Mimo V2.5 led at 1.4 trillion tokens per day, followed by DeepSeek V4 Flash (943.9B) and Tencent Hy3 (590B). Kimi K3 entered at #9 — the fastest new climber on the board.
No. The board sorts by token volume, which favors cheap high-throughput workloads. On hard reasoning, Claude Sonnet 4.6 and Opus 4.7 each hold 13.5% spend share. For stable agent hosting, see the pricing page.
Start with DeepSeek V4 Flash and GLM 5.2 for coding, escalate stuck steps to Claude Opus 5. OpenRouter works well as an R&D sandbox; production needs latency and compliance review for your region.
Deploy OpenRouter or LiteLLM routing on an always-on cloud Mac. Region and networking guidance lives in the help center; match daily or monthly rental to project length.