Июльский рейтинг: Xiaomi #1, Китай ~46% token share
OpenRouter сортирует по token volume — это ground truth «что dev'ы реально шлют», не «кто умнее в leaderboard». На 25.07 top-3 daily: Xiaomi Mimo V2.5 (1,4T/day), DeepSeek V4 Flash (943,9B/day), Tencent Hy3 (590B/day). 7 из top-10 — китайские модели; US держат только Nemotron 3 Ultra (NVIDIA), Claude и Gemini.
| # | Model | Vendor | Tokens/day | 30d cumulative |
|---|---|---|---|---|
| 1 | Mimo V2.5 | Xiaomi | 1,4T | 31,2T |
| 2 | DeepSeek V4 Flash | DeepSeek | 943,9B | 23,6T |
| 3 | Hy3 | Tencent | 590B | 23,4T |
| 4 | Nemotron 3 Ultra (free) | NVIDIA | 428,6B | 9T |
| 5 | DeepSeek V4 Pro | DeepSeek | 413,7B | 11,6T |
| 6 | GLM 5.2 | Z.ai | 316,7B | 13,3T |
| 7 | MiniMax M3 | MiniMax | 262,5B | 15,1T |
| 8 | Step 3.7 Flash | StepFun | 204,8B | 5,9T |
| 9 | Kimi K3 | Moonshot AI | 157,6B | 1,6T (new entry) |
| 10 | Ling 3.0 Flash | Ant InclusionAI | 128,3B | 417,3B |
| 11 | Gemini 3 Flash Preview | 106,3B | 4T | |
| 12 | Claude Sonnet 5 | Anthropic | 99,5B | 3,6T |
Vendor rollup: китайские labs суммарно ~46% token share — год назад <2%. US big three (OpenAI, Anthropic, Google) скатились с ~70% до 30–36%. Драйвер — price arbitrage: DeepSeek V4 Flash input ~$0,05–0,14/M vs GPT-5.5 ~$5/M, ~35× spread.
DeepSeek — самый стабильный vendor: ~16–18% share, но monthly champion model меняется.
Xiaomi — max volatility: Mimo V2.5 качает vendor share между 8% и 18%.
Kimi K3 — fastest riser: debut в top-9, 1,4TB open weights — record dump.
Daily noise: Claude Opus 4.8 был top-10 на 24.07, на 25.07 вытеснен Ling 3.0 Flash.
Snapshot ≠ trend: важнее retention в top tier, чем один lucky day.
Volume ≠ quality: dumbbell market structure
Reality check: OpenRouter rank = token throughput, not IQ. Дешёвые fast models за high-traffic apps легко залетают в top — даже если complex reasoning так себе.
| Axis | By token volume | By spend (hard tasks) |
|---|---|---|
| Leaders | Cheap CN open weights | Claude Sonnet 4.6 / Opus 4.7 по 13,5% |
| Typical workload | Chat, creative, roleplay, simple codegen | Complex reasoning, enterprise agents, classification |
| Price logic | Bulk, fault-tolerant, 35× delta | Closed flagship $5/$25 pricing power |
| July proof | Mimo V2.5 1,4T/day | Claude Opus 5 FrontierBench v0.1: 43,3% |
Рынок самосегментируется: cheap open weights жрут low-barrier tasks; closed flagship держат pricing power и security rep на hard workloads.
Spend by task type: general chat 35,7%, agent workflows 30,4%, code 26,5%, data processing 7,5%. Slice «classification/complex reasoning»: GPT-5.5 — 11,6%, #3; cheap models из volume chart тут почти invisible. Claude Opus 5 (drop 24.07) — 43,3% FrontierBench v0.1 (GPT-5.6 Sol 37,5%), pricing Opus tier $5/$25 per M.
| Model | Input/M | Output/M | Slot |
|---|---|---|---|
| DeepSeek V4 Flash | $0,05–0,14 | $0,24–0,28 | $/perf king, agentic coding default |
| GLM 5.2 | $0,45 | $3,31 | Opus-like planning в open weights |
| MiniMax M3 | $0,10 | $1,21 | Long context + multimodal budget pick |
| Kimi K3 | ~$3 | ~$15 | 1,4TB weights, closed-tier capability |
| Claude Opus 5 | $5 (fast $10) | $25 (fast $50) | Closed SOTA, июль benchmark #1 |
App layer: coding agents rule, roleplay — hidden half
Model chart = «какой brain популярен»; app layer (openrouter.ai/apps) = что на нём реально крутят.
| # | App | Type | Share (~) |
|---|---|---|---|
| 1 | Hermes Agent | Personal agent / CLI | ~45% |
| 2 | Kilo Code | Coding agent | ~13% |
| 3 | OpenClaw | Universal agent | ~9% |
| 4 | Claude Code | Coding agent | ~6% |
| 5 | Descript | Content pipeline | ~4,5% |
| 6–10 | pi, Lemonade, ISEKAI ZERO, Janitor AI, Cline | Agent / roleplay | 1,7%–3,3% |
Geek note: lineage Cline → Roo Code → Kilo Code — три форка одной code-agent ветки; «внук» Kilo Code уже обогнал предков. OpenRouter × a16z State of AI: creative roleplay >50% open-source volume — enterprise media это почти не освещает.
| Scenario | Pick | Why |
|---|---|---|
| Daily coding / agentic | DeepSeek V4 Flash | Best $/token, #2 по volume |
| Open planning quality | GLM 5.2 | Closest Opus-style planner в OSS |
| Hard reasoning | Claude Opus 5 / Sonnet 5 | Spend share leader на tough tasks |
| Long context OSS | Kimi K3 | 1M context, 1,4TB weights (pre-quant) |
| Multimodal budget | MiniMax M3 | Image input без разорения |
| US full OSS | Nemotron 3 Ultra | NVIDIA stack, free tier exists |
API routing: 6-step tiered runbook для prod
Task tiering: split bulk (chat/summary/simple completion) vs hard (reasoning/high-risk agent decisions). Single-model-for-all — anti-pattern.
Bulk tier default: DeepSeek V4 Flash — input $0,05–0,14/M; fallback GLM 5.2 когда нужен stronger open planning.
Hard tier: Claude Opus 5 reserved — escalate только после repeated cheap-tier failures; hybrid routing > full Opus burn.
OpenRouter unified endpoint — один key, сотни models, идеален для A/B; закладывайте ~180–250ms EU→US latency.
Security scorecard: sandbox escape incidents, vendor track record, least-privilege для agent tools; audit permissions перед prod.
7×24 host: agent gateway на cloud Mac, не на sleeping laptop; setup — центр помощи.
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-d '{
"models": [
"deepseek/deepseek-v4-flash",
"anthropic/claude-opus-5"
],
"messages": [{"role": "user", "content": "Plan a multi-step refactor..."}]
}'
Август 2026: forecast + citable hard numbers
Extrapolation из июльского тренда:
CN open weights → ~50% — unless US vendors slash prices (no signal yet).
Monthly champion keeps rotating — price war между Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, Moonshot не остановится.
Anthropic budget tier likely — Opus 5 = 4th flagship за 2 месяца; multi-price coverage strategy.
Kimi K3 1,4TB: community quant за 2–4 недели — до тех пор выигрывают inference clouds.
Security/compliance = new selection variable — US «AI Kill Switch» bill + frontier pre-release review framework ожидаются до августа.
CN vendor token share: ~46% (12 мес назад <2%) — один из steepest share migrations в AI.
Hermes Agent app share: ~45% — largest single app на OpenRouter.
DeepSeek vs GPT-5.5 input price: ~35× gap ($0,05–0,14/M vs ~$5/M).
Heads up: rankings shift daily — перед deploy сверяйте openrouter.ai/rankings.
July takeaway: capability и popularity расходятся. ROI выше в eval framework + tier routing, чем в погоне за #1 snapshot.
Multi-model agent gateway на laptop = sleep disconnects, RAM pressure, network jitter. Для 7×24 Hermes Agent, OpenClaw или multi-model CI: MESHLAUNCH Mac Mini cloud bare-metal — dedicated Apple Silicon, day/week/month billing. Цены: аренда.
На 25.07: Xiaomi Mimo V2.5 — 1,4T tokens/day, затем DeepSeek V4 Flash (943,9B) и Tencent Hy3 (590B). Kimi K3 — новый #9.
Нет — token volume ≠ intelligence. На hard tasks Claude Sonnet 4.6 и Opus 4.7 по 13,5% spend each. Stable agent host: цены аренды.
Coding: DeepSeek V4 Flash + GLM 5.2 first; Claude Opus 5 — только на stuck steps. OpenRouter = sandbox; prod — stable cloud + compliance review.
Deploy OpenRouter/LiteLLM router на cloud Mac. Region & config: центр помощи; day/month rental по длительности проекта.