Kimi K3: спецификация 2,8 трln open-source MoE
16.07.2026 Moonshot AI выпустил Kimi K3 — sparse MoE на 2,8 трln (2,8T) параметров, на ~75% больше прежнего рекордсмена DeepSeek V4 Pro (1,6T), в 2,7× больше Xiaomi open model (1,02T) и более чем в 7× больше Alibaba (397B). Конtrast: один tech blog, pricing page и сразу callable model ID kimi-k3 — без PR-шоу.
896 experts, 16 активируются на forward pass (sparsity 1,8%); 1 048 576 tokens (1M) context и native vision для сложного coding, long-document reasoning и knowledge work. Сводка: open-source, multimodal (text/image/video), ultra-long memory coding LLM — input на 40% дешевле Claude Opus 4.8; полные weights на Hugging Face 27.07.2026.
Scale shock: после взлёта DeepSeek доля Moonshot сжалась; K3 — стратегический ответ. За 12 месяцев серия Kimi 9 месяцев держала верхнюю границу scale среди open models.
Timing: релиз накануне WAIC 2026 (17–20.07); ожидаются дополнительные анонсы.
Commercial traction: к июню 2026 ARR >$300M, 6-й раунд в этом году, pre-money valuation $31,5B; >70% revenue — API; overseas paid users +400% YoY.
Long-context pain: full attention на 1M tokens даёт катастрофический рост KV cache; KDA создан именно под эту проблему.
Coding agent demand: SWE Marathon и подобные harness требуют часовых сессий с coherent context — 1M window + cache hit 90%+ как core selling point.
| Параметр | Значение |
|---|---|
| Total params | 2,8 трln (крупнейшая open-source модель) |
| Architecture | KDA + AttnRes + Stable LatentMoE |
| Active experts | 16 / 896 (sparse MoE, 1,8%) |
| Context window | 1 048 576 tokens (1M) |
| Input modalities | text, image, video |
| Reasoning mode | always-on max effort (low/high — позже) |
| Open weights | 27.07.2026, Hugging Face |
Kimi K3: KDA, AttnRes, Stable LatentMoE и бенчмарки
Kimi Delta Attention (KDA) — hybrid linear attention: слои чередуются 3:1 (3 linear + 1 full attention). Linear layers обрабатывают local structure дёшево; full attention сохраняет global signal. Итог: KV cache memory −75%, decode на 1M context до 6,3× быстрее; на short/long context и RL scaling обходит pure full-attention baseline.
Attention Residuals (AttnRes) перестраивает depth-wise residual: стандартный residual размывает ранние representations; AttnRes добавляет selective retrieval — модель напрямую подтягивает high-value features из ранних слоёв. ~25% training efficiency при overhead <2% compute.
Stable LatentMoE стабилизирует training при 896 experts / 16 active (extreme sparsity). Стек техник:
| Техника | Роль |
|---|---|
| Quantile Balancing | expert allocation из quantiles router scores — без heuristic hyperparams |
| Per-Head Muon | per-attention-head optimization для adaptive large-scale training |
| Sigmoid Tanh Unit (SiTU) | улучшенный activation control |
| Gated MLA | selective attention |
Суммарно K3 vs Kimi K2: scaling efficiency ~2,5×. Ниже self-reported benchmarks Moonshot (разные harness per vendor; independent replication in progress):
| Benchmark | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol | Claude Opus 4.8 | GLM-5.2 |
|---|---|---|---|---|---|
| DeepSWE | 67.5 | 70.0 | 73.0 | 59.0 | 46.2 |
| Program Bench | 77.8 | 76.8 | 77.6 | 71.9 | 63.7 |
| Terminal Bench 2.1 | 88.3 | 84.6 | 88.8 | 84.6 | 82.7 |
| FrontierSWE | 81.2 | 86.6 | 71.3 | 66.7 | 67.3 |
| SWE Marathon | 42.0 | 35.0 | 39.0 | 40.0 | 13.0 |
| BrowseComp | 91.2 | 88.0 | 90.4 | 84.3 | — |
| Automation Bench | 30.8 | 29.1 | 29.7 | 27.2 | 12.9 |
| GPQA-Diamond | 93.5 | 92.6 | 94.1 | 91.0 | 91.2 |
| MMMU-Pro (vision) | 81.6 | 81.2 | 83.0 | 78.9 | — |
| OmniDocBench | 91.1 | 89.8 | 85.8 | 87.9 | — |
SWE Marathon тестирует sustained long-horizon coding — K3 лидирует с 42.0. Artificial Analysis Intelligence Index v4.1: K3 — 57.1 (#4 overall), за Claude Fable 5 (59.9) и GPT-5.6 Sol (58.9) с gap 2.8 points; highest среди open-weight models.
Kimi K3 pricing: сравнение с Claude, GPT и DeepSeek
Standard API: $3/M input, $15/M output — паритет с Claude Sonnet 5, но 5× context (1M vs 200K). Cache hit: $0,30/M (1/10 standard); Moonshot reports >90% cache hit в coding workflows — effective input cost минимален. OpenRouter 7-day weighted average: ~$0,55/M effective input.
| Model | Input ($/M) | Output ($/M) | Cache hit input | Context |
|---|---|---|---|---|
| Kimi K3 | $3.00 | $15.00 | $0.30 | 1M |
| Claude Sonnet 5 | $3.00 (promo $2) | $15.00 (promo $10) | — | 200K |
| Claude Opus 4.8 | $5.00 | $25.00 | — | 200K |
| GPT-5.5 | $5.00 | $30.00 | — | 400K |
| DeepSeek V4 Pro | $1.74 | $3.48 | $0.145 | 128K |
| Kimi K2.6 | $0.95 | $4.00 | $0.16 | 256K |
Vs Claude Opus 4.8: K3 сильнее на ряде benchmarks при 60% input и 40% output cost. China API: ¥20/M input, ¥100/M output, ¥2/M cache hit. Consumer tier на kimi.com — free account с K3 max effort; prepaid от ¥199 (promo до 11.08). Бюджет инфраструктуры: цены аренды Mac mini.
Kimi K3: четыре способа подключения и шестишаговый runbook
Kimi web/app: kimi.com, регистрация (Google SSO), K3 в max effort по умолчанию, без credit card.
API key: platform.kimi.ai — developer account, create key, top-up balance.
OpenAI-compatible client: base_url="https://api.moonshot.ai/v1", model kimi-k3.
OpenRouter: model ID moonshotai/kimi-k3, official Moonshot pricing без markup, full 1M context.
Prompt caching: routing cache key в coding pipeline — leverage 90%+ hit rate, effective input ~$0,55/M.
27.07 milestone: full weights на Hugging Face; MXFP4/NVFP4 quant + Day-0 vLLM/SGLang expected; production — supernode 64+ GPUs.
from openai import OpenAI
client = OpenAI(
api_key="your_moonshot_api_key",
base_url="https://api.moonshot.ai/v1"
)
response = client.chat.completions.create(
model="kimi-k3",
messages=[{"role": "user", "content": "Разбери этот фрагмент кода..."}]
)
Open-source timeline: WAIC 17–20.07 — доп. анонсы → 27.07.2026 full K3 weights — первый downloadable open model >2 трln params.
Kimi K3: матрица сценариев и cite-ready numbers
| Сценарий | Модель | Почему |
|---|---|---|
| Sustained long coding | Kimi K3 | SWE Marathon #1, longest 1M context |
| Complex repo-level bugfix | Claude Fable 5 | FrontierSWE 86.6 — large lead |
| Terminal/toolchain agent | GPT-5.6 Sol | Terminal Bench 2.1 + Coding Agent Index |
| Ultra-long doc / multimodal | Kimi K3 | OmniDocBench 91.1, native vision + 1M |
| Cost-sensitive | DeepSeek V4 Pro | output $3,48/M vs $15/M у K3 |
| Self-host (post 27.07) | Kimi K3 | strongest open weights, MXFP4-friendly |
Param scale: 2,8T — +~75% vs DeepSeek V4 Pro (1,6T); largest downloadable open model после 27.07.
KDA efficiency: KV cache −75%, 1M decode +6,3×, scaling efficiency +2,5× vs K2.
Intelligence Index: Artificial Analysis v4.1 — 57.1, top open model, 2.8 points от closed flagship tier.
Caveat: benchmarks — self-reported Moonshot; K3 через Kimi Code, GPT через Codex, Claude через Claude Code — разные harness. Independent third-party replication ongoing; трактуйте как directional signal.
Kimi K3 — не param vanity: KDA, AttnRes и Stable LatentMoE — инженерные решения с measurable gains на long-horizon coding и document understanding, на части метрик паритет или lead vs closed flagship. Pure API — быстрый старт, но локальный Mac под Kimi Code agent loops часто упирается в memory pressure и IDE overhead; cloud VM теряет на Metal compatibility. Командам с 7×24 coding agents, parallel dev servers и persistent inference bare-metal Mac mini аренда у MESHLAUNCH — типичный production path: dedicated Apple Silicon, elastic day/week/month billing, iOS CI/CD и agent automation без shared-tenant surprises. Заказ: цены аренды, onboarding — центр помощи.
Sources: Moonshot AI official blog · Kimi API Platform docs · Artificial Analysis · OpenRouter pricing · VentureBeat · SCMP
Да — kimi.com, free account, K3 в max effort. API: $3/$15 per 1M tokens. Бюджет инфра: цены аренды.
Full weights — 27.07.2026 на Hugging Face. Production: supernode 64+ accelerator cards; consumer local deploy нереалистичен. Training: MXFP4 weights + MXFP8 activations — quant-friendly.
K3: ~2× params, 8× context, сильнее coding/doc benchmarks. DeepSeek V4 Pro: output $3,48/M — win на cost-sensitive workloads. Детали: центр помощи.
Да — whole-repo analysis, long legal/research docs, multi-turn agent memory. Flat pricing без length surcharge; 90%+ cache hit держит effective cost низким.
Moonshot AI: low и high effort — в последующих updates. Сейчас только max effort always-on.