Kimi K3: обзор 2026
2,8 трln параметров — новый рекорд open-source

KDA · 1M context · benchmarks · pricing · API · open weights 27.07

Kimi K3 2,8 трln open-source LLM обзор 2026
В ночь на 16.07.2026 Moonshot AI (月之暗面) тихо повесил в API-документации «Kimi K3 is live» — без keynote, но с крупнейшей по параметрам open-source моделью в мире (2,8 трln). Если вы выбираете coding agent для команды, сравниваете с Claude Fable 5 и GPT-5.6 Sol или ждёте downloadable weights, этот разбор закрывает: ① что решают KDA / AttnRes / Stable LatentMoE; ② как K3 выступает на SWE Marathon, OmniDocBench и смежных бенчмарках; ③ $3/$15 pricing, четыре способа подключения и что означает open weights 27.07.2026.
01

Kimi K3: спецификация 2,8 трln open-source MoE

16.07.2026 Moonshot AI выпустил Kimi K3 — sparse MoE на 2,8 трln (2,8T) параметров, на ~75% больше прежнего рекордсмена DeepSeek V4 Pro (1,6T), в 2,7× больше Xiaomi open model (1,02T) и более чем в 7× больше Alibaba (397B). Конtrast: один tech blog, pricing page и сразу callable model ID kimi-k3 — без PR-шоу.

896 experts, 16 активируются на forward pass (sparsity 1,8%); 1 048 576 tokens (1M) context и native vision для сложного coding, long-document reasoning и knowledge work. Сводка: open-source, multimodal (text/image/video), ultra-long memory coding LLM — input на 40% дешевле Claude Opus 4.8; полные weights на Hugging Face 27.07.2026.

01

Scale shock: после взлёта DeepSeek доля Moonshot сжалась; K3 — стратегический ответ. За 12 месяцев серия Kimi 9 месяцев держала верхнюю границу scale среди open models.

02

Timing: релиз накануне WAIC 2026 (17–20.07); ожидаются дополнительные анонсы.

03

Commercial traction: к июню 2026 ARR >$300M, 6-й раунд в этом году, pre-money valuation $31,5B; >70% revenue — API; overseas paid users +400% YoY.

04

Long-context pain: full attention на 1M tokens даёт катастрофический рост KV cache; KDA создан именно под эту проблему.

05

Coding agent demand: SWE Marathon и подобные harness требуют часовых сессий с coherent context — 1M window + cache hit 90%+ как core selling point.

ПараметрЗначение
Total params2,8 трln (крупнейшая open-source модель)
ArchitectureKDA + AttnRes + Stable LatentMoE
Active experts16 / 896 (sparse MoE, 1,8%)
Context window1 048 576 tokens (1M)
Input modalitiestext, image, video
Reasoning modealways-on max effort (low/high — позже)
Open weights27.07.2026, Hugging Face
02

Kimi K3: KDA, AttnRes, Stable LatentMoE и бенчмарки

Kimi Delta Attention (KDA) — hybrid linear attention: слои чередуются 3:1 (3 linear + 1 full attention). Linear layers обрабатывают local structure дёшево; full attention сохраняет global signal. Итог: KV cache memory −75%, decode на 1M context до 6,3× быстрее; на short/long context и RL scaling обходит pure full-attention baseline.

Attention Residuals (AttnRes) перестраивает depth-wise residual: стандартный residual размывает ранние representations; AttnRes добавляет selective retrieval — модель напрямую подтягивает high-value features из ранних слоёв. ~25% training efficiency при overhead <2% compute.

Stable LatentMoE стабилизирует training при 896 experts / 16 active (extreme sparsity). Стек техник:

ТехникаРоль
Quantile Balancingexpert allocation из quantiles router scores — без heuristic hyperparams
Per-Head Muonper-attention-head optimization для adaptive large-scale training
Sigmoid Tanh Unit (SiTU)улучшенный activation control
Gated MLAselective attention

Суммарно K3 vs Kimi K2: scaling efficiency ~2,5×. Ниже self-reported benchmarks Moonshot (разные harness per vendor; independent replication in progress):

BenchmarkKimi K3Claude Fable 5GPT-5.6 SolClaude Opus 4.8GLM-5.2
DeepSWE67.570.073.059.046.2
Program Bench77.876.877.671.963.7
Terminal Bench 2.188.384.688.884.682.7
FrontierSWE81.286.671.366.767.3
SWE Marathon42.035.039.040.013.0
BrowseComp91.288.090.484.3
Automation Bench30.829.129.727.212.9
GPQA-Diamond93.592.694.191.091.2
MMMU-Pro (vision)81.681.283.078.9
OmniDocBench91.189.885.887.9

SWE Marathon тестирует sustained long-horizon coding — K3 лидирует с 42.0. Artificial Analysis Intelligence Index v4.1: K3 — 57.1 (#4 overall), за Claude Fable 5 (59.9) и GPT-5.6 Sol (58.9) с gap 2.8 points; highest среди open-weight models.

03

Kimi K3 pricing: сравнение с Claude, GPT и DeepSeek

Standard API: $3/M input, $15/M output — паритет с Claude Sonnet 5, но 5× context (1M vs 200K). Cache hit: $0,30/M (1/10 standard); Moonshot reports >90% cache hit в coding workflows — effective input cost минимален. OpenRouter 7-day weighted average: ~$0,55/M effective input.

ModelInput ($/M)Output ($/M)Cache hit inputContext
Kimi K3$3.00$15.00$0.301M
Claude Sonnet 5$3.00 (promo $2)$15.00 (promo $10)200K
Claude Opus 4.8$5.00$25.00200K
GPT-5.5$5.00$30.00400K
DeepSeek V4 Pro$1.74$3.48$0.145128K
Kimi K2.6$0.95$4.00$0.16256K

Vs Claude Opus 4.8: K3 сильнее на ряде benchmarks при 60% input и 40% output cost. China API: ¥20/M input, ¥100/M output, ¥2/M cache hit. Consumer tier на kimi.com — free account с K3 max effort; prepaid от ¥199 (promo до 11.08). Бюджет инфраструктуры: цены аренды Mac mini.

04

Kimi K3: четыре способа подключения и шестишаговый runbook

01

Kimi web/app: kimi.com, регистрация (Google SSO), K3 в max effort по умолчанию, без credit card.

02

API key: platform.kimi.ai — developer account, create key, top-up balance.

03

OpenAI-compatible client: base_url="https://api.moonshot.ai/v1", model kimi-k3.

04

OpenRouter: model ID moonshotai/kimi-k3, official Moonshot pricing без markup, full 1M context.

05

Prompt caching: routing cache key в coding pipeline — leverage 90%+ hit rate, effective input ~$0,55/M.

06

27.07 milestone: full weights на Hugging Face; MXFP4/NVFP4 quant + Day-0 vLLM/SGLang expected; production — supernode 64+ GPUs.

Python
from openai import OpenAI

client = OpenAI(
    api_key="your_moonshot_api_key",
    base_url="https://api.moonshot.ai/v1"
)

response = client.chat.completions.create(
    model="kimi-k3",
    messages=[{"role": "user", "content": "Разбери этот фрагмент кода..."}]
)

Open-source timeline: WAIC 17–20.07 — доп. анонсы → 27.07.2026 full K3 weights — первый downloadable open model >2 трln params.

05

Kimi K3: матрица сценариев и cite-ready numbers

СценарийМодельПочему
Sustained long codingKimi K3SWE Marathon #1, longest 1M context
Complex repo-level bugfixClaude Fable 5FrontierSWE 86.6 — large lead
Terminal/toolchain agentGPT-5.6 SolTerminal Bench 2.1 + Coding Agent Index
Ultra-long doc / multimodalKimi K3OmniDocBench 91.1, native vision + 1M
Cost-sensitiveDeepSeek V4 Prooutput $3,48/M vs $15/M у K3
Self-host (post 27.07)Kimi K3strongest open weights, MXFP4-friendly
A

Param scale: 2,8T — +~75% vs DeepSeek V4 Pro (1,6T); largest downloadable open model после 27.07.

B

KDA efficiency: KV cache −75%, 1M decode +6,3×, scaling efficiency +2,5× vs K2.

C

Intelligence Index: Artificial Analysis v4.1 — 57.1, top open model, 2.8 points от closed flagship tier.

Caveat: benchmarks — self-reported Moonshot; K3 через Kimi Code, GPT через Codex, Claude через Claude Code — разные harness. Independent third-party replication ongoing; трактуйте как directional signal.

Kimi K3 — не param vanity: KDA, AttnRes и Stable LatentMoE — инженерные решения с measurable gains на long-horizon coding и document understanding, на части метрик паритет или lead vs closed flagship. Pure API — быстрый старт, но локальный Mac под Kimi Code agent loops часто упирается в memory pressure и IDE overhead; cloud VM теряет на Metal compatibility. Командам с 7×24 coding agents, parallel dev servers и persistent inference bare-metal Mac mini аренда у MESHLAUNCH — типичный production path: dedicated Apple Silicon, elastic day/week/month billing, iOS CI/CD и agent automation без shared-tenant surprises. Заказ: цены аренды, onboarding — центр помощи.

Sources: Moonshot AI official blog · Kimi API Platform docs · Artificial Analysis · OpenRouter pricing · VentureBeat · SCMP

FAQ

Да — kimi.com, free account, K3 в max effort. API: $3/$15 per 1M tokens. Бюджет инфра: цены аренды.

Full weights — 27.07.2026 на Hugging Face. Production: supernode 64+ accelerator cards; consumer local deploy нереалистичен. Training: MXFP4 weights + MXFP8 activations — quant-friendly.

K3: ~2× params, 8× context, сильнее coding/doc benchmarks. DeepSeek V4 Pro: output $3,48/M — win на cost-sensitive workloads. Детали: центр помощи.

Да — whole-repo analysis, long legal/research docs, multi-turn agent memory. Flat pricing без length surcharge; 90%+ cache hit держит effective cost низким.

Moonshot AI: low и high effort — в последующих updates. Сейчас только max effort always-on.