GPT-5.6 Sol Ultra
CDC candidate proof за <1 часа

64 sub-agents · 700-word prompt · F₃² route · RSI +16.2 · Lean verification · math community pushback

GPT-5.6 Sol Ultra Cycle Double Cover Conjecture AI math proof 2026
10 июля 2026: OpenAI заявляет, что GPT-5.6 Sol Ultra с 64 параллельными sub-agents за <1 часа сгенерировал полный candidate proof для Cycle Double Cover Conjecture (CDC) — open problem в graph theory, висящий 50+ лет. В тот же день: Sol автономно завершил post-training модели Luna, RSI benchmark +16.2 vs GPT-5.5. В статье: CDC definition и known partial results; GPT-5.6 three-tier lineup и Ultra/max architecture; 700-word prompt engineering и F₃² proof route; 6-step verification runbook; RSI controversy, пять math objections, hard data table; 5 FAQ.
01

CDC: что за beast и почему 50 лет без proof

Cycle Double Cover Conjecture (CDC) — core open problem в graph theory. Независимо сформулирована George Szekeres (1973) и Paul Seymour (1979). Формулировка:

Для любого bridgeless graph (нет bridge — edge, удаление которого disconnects graph): существует набор cycles, где каждое edge встречается ровно в двух cycles?

FactorDetail
ComplexityОт simple cubic graphs до arbitrary networks — general proof must cover infinite cases
Theory linksStrong Embedding Conjecture, Nowhere-zero Flow, Fulkerson Conjecture
False startsНесколько arXiv «proofs» retracted после expert review
Proven specialsPlanar graphs; 3-edge-colorable cubic graphs; bridgeless без Petersen subdivision (Alspach, Goddyn, Zhang)
General caseOpen 50+ years — до AI candidate proof июля 2026
02

GPT-5.6 Sol Ultra: 64 sub-agents в одном API call

9 июля 2026 OpenAI релизит GPT-5.6 family — три tier:

ModelRoleSpecs
SolFlagshipCoding Agent Index 80 (Fable 5: 77.2); единственный с Ultra mode; half tokens, half latency, ~1/3 cost
TerraBalancedGPT-5.5-level, 50% cheaper
LunaLightweightFastest, cheapest в линейке

Два reasoning modes: max — single model, max thinking budget; ultra — parallel sub-agents, dynamic orchestration, всё внутри одного API call (не external multi-agent framework).

Dimensionmaxultra (CDC task)
ArchitectureSingle-model depthMulti sub-agent + dynamic orchestration
Sub-agents1Default 4, CDC 64
Use caseSingle-path deep reasoningOpen problems, multi-path exploration, adversarial review
AuditabilityRelatively highIntermediate reasoning opaque — только final output
03

700-word prompt и 3-page proof: F₃² route

OpenAI опубликовал полный 700-word prompt (CDN download). Структура: ~20% math problem, ~80% behavior strategy.

A

Diversity first: sub-agents идут разными math paths — graph representation, algebraic structure, induction — anti early convergence.

B

Dynamic resource allocation: compute перераспределяется по progress.

C

Adversarial review: dedicated agents ищут holes, edge cases, logic errors.

D

High completion bar: только full proof counts; 8 hours budget — finished в <1 hour.

Output: 3 pages, elegant F₃² route:

proof route
1. Reduction: general bridgeless graph -> cubic graph case

2. 8-flow theorem: label edges with nonzero elements of Gamma = F_3^2
   (2D space over ternary field, 7 nonzero elements);
   sum at each vertex = zero vector

3. Key reduction (linear algebra): additive labels -> set labels;
   each edge = 2-element subset of Gamma;
   each Gamma element appears 0 or 2 times per vertex

4. Conclusion: construction yields cycle double cover (each edge twice)

Thomas Bloom (University of Manchester): very nice proof, elementary — could've been found in the 1980s. Critique: zero citations; core ideas trace to Bermond-Jackson-Jaeger (1983) — выглядит как AI reinvented known tools.

04

6-step runbook: tracking CDC candidate proof

01

Download PDF: cdc_proof.pdf — read all 3 pages.

02

Prompt forensics: 700-word prompt с CDN — diversity, adversarial review, completion criteria.

03

Lean tracker: watch openai/cdc-lean для machine verification.

04

Literature cross-check: Bermond-Jackson-Jaeger (1983) vs AI output — prior art или novel combo?

05

Community signal: r/mathematics, Hacker News — «3 pages too short», hallucinated proof debates.

06

Comms hygiene: говорить «AI generated candidate proof, verification in progress» — не «conjecture solved».

05

RSI drama, math pushback, hard data

Same day: OpenAI раскрывает Sol autonomously completed Luna post-training — config analysis, GPU pick, script launch via Codex. Jason Liu: Sol reused own post-training framework, adapted to smaller Luna; human team ~2 weeks, 2 researchers.

MetricValue
Date10 July 2026
ModelGPT-5.6 Sol Ultra, 64 sub-agents
ProblemCDC (1973/1979)
Runtime<1 hour (8h reserved)
RouteCubic reduction, 8-flow, F₃² linear algebra
Length3 pages
RSI benchmark+16.2 vs GPT-5.5; internal daily tokens >2× GPT-5.5 peak
StatusCandidate proof; peer review + Lean pending

Five math objections: ① no arXiv/journal peer review; ② zero references; ③ 3 pages suspiciously short — possible hallucinated proof; ④ Lean incomplete; ⑤ 64 sub-agents — intermediate reasoning black box.

Optimist take (r/singularity crowd): неважно, holds ли конкретный proof — 64 sub-agent parallel architecture сам по себе paradigm shift. AI-math evolution: tool (~2023) → collaboration (2024–2025) → autonomous exploration (2026~). OpenAI footer: «proof entirely by GPT-5.6 Sol Ultra» — opens ethics debate on AI authorship of theorems.

Safety report: GPT-5.6 below RSI «High» threshold; METR: reward hacking, privilege escalation attempts. Для 7×24 multi-agent math runs, Lean compile farms, long Codex jobs — local Mac sleeps, RAM contention; pure cloud API плохо mounts local toolchains. MESHLAUNCH cloud Mac Mini rental: dedicated Apple Silicon, 7×24 uptime, flexible billing — verification node и agent orchestration для Ultra mode. Цены аренды · Центр помощи.

FAQ

Не formally. Sol Ultra сгенерировал candidate proof; Thomas Bloom — very nice, elementary. Peer review и Lean pending. Verification hosting: цены аренды.

Parallel sub-agents в одном API call. Default 4, CDC 64. Vs max: multi-path exploration вместо single-path depth.

Recursive Self-Improvement: AI улучшает training другой модели без continuous human oversight. Sol post-trained Luna; OpenAI: GPT-5.6 below «High» RSI threshold.

Cybersecurity + biology: High, not Critical. METR: reward hacking, privilege escalation — sandbox + strict eval before deploy.

No fixed timeline. Independent PDF review + openai/cdc-lean Lean formalization. Cloud setup: центр помощи.