Is DeepSeek Building Its Own AI Chip?
Inside the July 2026 Reuters Report

Unit economics · $7.4B funding · T-Head mass production · Global custom silicon wave · Five drivers

DeepSeek custom AI chip and Alibaba T-Head silicon progress
In June 2026, OpenAI and Broadcom taped out Jalapeño, a custom inference ASIC. One week later, Reuters reported that DeepSeek is developing its own inference chip—while already running on Huawei Ascend. This is not a China-only story; it is a unit economics story. For developers and infra leads, this guide covers: ① the global custom-silicon wave; ② what Reuters confirmed vs. rumor; ③ Liang Wenfeng's past remarks; ④ Alibaba T-Head mass production; ⑤ five drivers and a six-step runbook. Last updated: July 9, 2026. DeepSeek has not officially confirmed the chip project as of this writing.
01

This Isn't Just China: OpenAI Jalapeño and the Global Custom Chip Wave

By July 2026, "AI companies building chips" is a global phenomenon. TrendForce data shows hyperscaler custom AI chip shipment growth at 44.6%, far outpacing general-purpose GPUs at 16.1%—custom silicon is winning on growth for the first time.

01

2026-06-24: OpenAI + Broadcom announce Jalapeño inference ASIC—9 months from design to tape-out.

02

2026-07-02: The Information reports Anthropic in talks with Samsung on 2nm custom silicon.

03

2026-07-07: Reuters: DeepSeek developing custom inference chip (exclusive).

04

2026-07-07: The Information: Zhipu AI evaluating custom chips.

05

Core thesis: AI competition has shifted from "best model" to "cheapest, most controllable compute."

Training is like a down payment; inference is rent that scales with daily active users. At ChatGPT-scale usage, inference spend exceeds training. Custom ASICs target that recurring bill.

02

What Reuters Actually Reported (And What DeepSeek Hasn't Confirmed)

On July 7–8, 2026, outlets followed a Reuters exclusive with consistent details: DeepSeek is building a custom AI chip for inference only, not training. The project reportedly started around mid-2025 and remains early-stage. The company is talking to chip designers, foundries, and memory suppliers, and has stepped up private hiring of chip engineers—not via public job boards.

Counterintuitive but important: DeepSeek already adapts to Huawei Ascend (V4 in April 2026; V4-Flash partially trained on Ascend) yet still pursues in-house silicon. The accurate framing is partnership and self-development in parallel—partnership is live, self-development is early. Success would reduce dual dependence on Nvidia and Huawei Ascend.

DimensionAssessment
Source qualityHigh—Reuters "three people familiar with the matter"
Official confirmationNone as of July 9, 2026
Circumstantial evidenceStrong—$7.4B round (June 2026) citing chip R&D; quiet hiring; UE8M0 FP8 format seen as hardware-software co-design signal
Contradictory takesSome analysts said DeepSeek would lean on Huawei short-term; reality is parallel tracks
Timeline
2023–2024  Liang Wenfeng interviews: export controls, compute hunger
2025-01    DeepSeek R1 on Nvidia H800 (already export-restricted)
Mid-2025   Reported in-house chip project start
2026-04    DeepSeek V4 on Huawei Ascend
2026-06    ~$7.4B external funding; chip R&D in disclosed use of proceeds
2026-07-07 Reuters: DeepSeek custom inference chip (exclusive)
03

What DeepSeek CEO Liang Wenfeng Has Said About Chips and Compute

Liang Wenfeng (DeepSeek CEO) rarely speaks publicly. The most useful sources are two deep interviews with "暗涌 Waves" in May 2023 and July 2024. He never announced a chip program. Reuters describes company actions—hiring, supplier talks—not a founder launch event.

"Our real challenge has never been funding—it is export controls on advanced chips." — Liang Wenfeng, July 2024

01

4× compute gap: Domestic best vs. abroad: ~2× training efficiency gap plus ~2× data efficiency gap—roughly 4× compute needed for parity.

02

Missing tech community: Domestic chips stall without ecosystem; someone must stand at the frontier with first-hand knowledge.

03

Endless compute appetite: Researchers always want more; DeepSeek deliberately deploys as much compute as possible.

These quotes establish strategic motive—constraints, export controls, co-design necessity—but are not a product announcement.

04

Alibaba's T-Head Is Already Shipping — Jack Ma's 2018 Bet Pays Off in 2026

Alibaba chip work is not a fresh rumor—it is an eight-year executed strategy. Do not write "Jack Ma recently said they would build chips." The accurate arc: Jack Ma set T-Head strategy in 2018; Joe Tsai explained export-control pressure in 2024; CEO Wu Yongming disclosed mass-production metrics in 2026.

At the September 2018 Cloud栖 conference, Alibaba merged Zhongtian Micro and Damo Academy chip teams into T-Head Semiconductor. Jack Ma personally named the unit. Chips became a group-level strategic priority.

FigureRolePublic chip-related stance
Jack Ma2018 strategic sponsorNamed T-Head; elevated chips to group strategy
Joe TsaiChairman2024 podcast: export limits clearly hit Alibaba Cloud; long-term belief in domestic advanced semiconductors
Wu YongmingCEOFY2026 earnings call: 470K+ AI chips delivered; billion-yuan annualized revenue; IPO option for T-Head
ModelTimingHighlights
GuangNeng 8002019Early AI inference chip
Zhenwu 810EJan 2026Train+infer; 96GB HBM2e; between Nvidia A800 and H20; in mass production
Zhenwu M8902026144GB; 800GB/s die-to-die; ~3× 810E
Zhenwu V900Planned Q3 2027216GB; 1200GB/s interconnect

Commercial data (H1 2026): 560,000+ units shipped; billion-yuan annualized revenue; 400+ enterprise customers on Zhenwu clusters. T-Head registered capital rose to 1 billion yuan (June 2026). Alibaba pledged 380 billion yuan over three years for cloud and AI infrastructure. WSJ: new chips compatible with Nvidia CUDA to ease migration; manufacturing shifting from early TSMC to domestic foundries (industry points to SMIC 7nm-class flows).

05

Why Tech Giants Build Custom AI Chips: Cost, Control, and the Nvidia Tax

Economics comes first. Custom inference ASICs can deliver 30–65% total cost of ownership (TCO) advantage at hyperscaler scale; per-token costs may fall 30–40%. Nvidia datacenter GPU gross margins exceed 70%—every H200 purchase ships most profit to Nvidia. In-house silicon converts recurring "GPU tax" into upfront R&D.

1

Economics: Morgan Stanley once estimated 24,000 Blackwell GPUs at ~$852M hardware vs. ~$99M for an equivalent Google TPU cluster (hardware-only, Breakingviews/Reuters).

2

Supply chain resilience: Export controls, allocation risk, single-vendor dependency—not just cybersecurity, but predictable supply.

3

Hardware-software co-design: DeepSeek UE8M0 FP8, OpenAI Jalapeño around real serving (KV cache, batching, latency). GPUs trade efficiency for flexibility; ASICs do the reverse for known workloads.

4

Bargaining power: Even partial self-supply strengthens Nvidia negotiations and full-stack storytelling (model + cloud + silicon).

5

Energy: Performance-per-watt matters at megawatt and gigawatt datacenters; ASICs strip unused GPU circuits.

CompanyProjectStageFocusKey metric
DeepSeekUnnamed inference ASICEarly R&DInference$7.4B funding; quiet hiring; unconfirmed
Alibaba T-HeadZhenwu 810E / M890Mass productionTrain+infer560K+ shipped; billion-yuan revenue
HuaweiAscend 950 seriesMass productionTrain+inferDeepSeek V4 adapted; orders surging
OpenAIJalapeño (Broadcom)Taped outInference9-month design cycle; deploy late 2026
GoogleTPU v6/v7At scaleTrain+inferGemini end-to-end on TPU
AmazonTrainium3 / InferentiaCommercialBothAnthropic on Trainium at scale
DimensionTrainingInference
WorkloadDynamic, experimental, frequent architecture changesStatic model, predictable request patterns
Software moatCUDA stack (cuDNN, NCCL, Nsight) very deepHand-tuned kernels for fixed models feasible
EconomicsLarge one-time cluster capex24/7 recurring spend—often larger at scale
VerdictTraining remains Nvidia territory; inference is the ASIC battleground

Six-step runbook for developers:

01

Label rumor vs. confirmed: Do not put Reuters-sourced early projects into procurement SLAs.

02

Split training vs. inference budgets: Most custom chips are inference-first; keep training on CUDA-class GPUs.

03

Benchmark with TCO, not list price: Model token cost, power, and multi-year depreciation together.

04

Track global peers: Jalapeño, Anthropic-Samsung talks, Zhipu evaluation—custom silicon growth is structural.

05

Plan supply-chain scenarios: Export controls can accelerate economics you already had.

06

Keep edge dev stable: Cheaper cloud inference does not replace local Agent/Xcode workflows—refresh this article every 2–4 weeks.

Risk: Early programs fail—Meta MTIA rebooted. Architecture shifts can obsolete ASICs. Write "reportedly" until official confirmation.

A

$7.4B DeepSeek round (June 2026): Disclosed uses include custom AI chips and domestic compute centers.

B

560K+ T-Head shipments: Wu Yongming FY2026 call; billion-yuan annualized revenue.

C

44.6% vs 16.1% growth: Custom AI silicon outpacing general GPUs (TrendForce 2026).

Buying a Mac for local Agent work looks cheaper when inference prices fall, but you still face memory ceilings, unstable 24/7 uptime, and build-queue contention. For production iOS CI/CD and AI Agent automation, MESHLAUNCH Mac Mini cloud rental is usually the better fit: dedicated bare metal, six regions, flexible daily/weekly/monthly terms.

FAQ

According to a July 7, 2026 Reuters report citing three sources, DeepSeek is in the early stages of developing a custom chip for AI inference. The company has not officially confirmed the project.

No public announcement. In 2024 interviews he said export controls on advanced chips were DeepSeek's main challenge, not funding—establishing motive, not a product launch.

Alibaba's chip unit T-Head (founded 2018) mass-produces Zhenwu AI chips—560,000+ units shipped, billion-yuan annualized revenue in 2026. See our pricing page for edge dev environments.

Inference workloads are repetitive and predictable—ideal for custom ASICs. Training still relies heavily on Nvidia GPUs and the CUDA software stack.

Both. Economics is the primary driver—cutting the Nvidia tax and per-token costs at scale—while export controls accelerate the shift. Check the help center for local dev setup guidance.