Section 09
Forward View
The next chapter turns on one question: can DeepSeek keep matching the frontier under a hard compute ceiling?
5 sources1 Chinese-languageAs of 2 Jun 2026
The R2 saga crystallized DeepSeek's central tension. A standalone R2 (targeted ~May 2025) slipped — reportedly on both Liang's perfectionism and chip shortages after April-2025 H20 curbs — and an attempt to train on Huawei Ascend reportedly failed. DeepSeek instead shipped V4 (Apr 2026), co-engineered for Ascend but trailing the frontier ~3–6 months. Talent and ideas are not the constraint; compute is.
The R2 / Huawei episode
The Information reported R2's delay reflected CEO Liang being unsatisfied with performance, compounded by a China-side Nvidia shortage after new H20 export curbs[64]. FT-sourced reporting added that Chinese authorities urged DeepSeek to train on Huawei Ascend, but persistent instability, slow interconnects and immature CANN software meant it never completed a fully successful Ascend training run — forcing a revert to Nvidia for training while using Ascend for inference[65]. Against that backdrop, Chinese press reported DeepSeek may pursue its first-ever external financing — read by some as building a compute war chest for the next generation[77].
What V4 tells us
DeepSeek effectively skipped a standalone R2 and previewed V4 on 24 April 2026 — V4-Pro (1.6T params, 1M-token context) co-engineered for Huawei chips — with reporting that it "falls marginally short of GPT-5.4 and Gemini 3.1 Pro"[66]. It continues to lead on price, with another permanent ~75% cut defending developer mindshare[76]. CSIS frames the stakes bluntly: the team would be formidable with more compute, but each Ascend 910C is only ~60% of an H100 for inference and its software is "difficult and unstable"[67].
Three scenarios to weigh
Not predictions — conditions to watch. Which one plays out hinges almost entirely on the compute question.
Bull
Domestic compute clicks
DeepSeek makes Huawei Ascend (or Cambricon) work at training scale, escaping the export-control ceiling. Its lean team + open ecosystem compound, and it re-takes the efficiency frontier on home silicon.[75]
Base
A strong fast-follower
DeepSeek stays ~3–6 months behind the frontier, leading on cost-efficiency and open weights but not dominating its home market. V4-class models keep it relevant via aggressive pricing.[66]
Bear
The compute ceiling bites
Constrained chips and an unstable Ascend stack cap the next scale-ups; talent keeps leaving; consumer usage keeps fading behind Doubao. DeepSeek becomes one capable lab among many.[67]
Weighed together, the three scenarios are not equally likely. The base case — a strong, cheap fast-follower — is the one the current evidence supports: the bull case requires an Ascend training breakthrough that has demonstrably not happened yet[65], while the bear case requires the price leadership and the release cadence to fail at the same time, and neither has[76].
The weighing
On whether the model was really built for ~$5.6M: the evidence leans true but narrow — a genuine final-run cost sitting on a far larger hardware base (high confidence). The controlling evidence is DeepSeek's own caveat that the figure excludes prior research and ablations[22] and SemiAnalysis's ~$1.6B server-CapEx estimate[23], which outweighs the "fabricated number" counter because the two figures come from opposing narratives yet are arithmetically compatible — they measure different things. The strongest surviving counter-argument: the ~50,000-Hopper-GPU estimate, which if confirmed would mean the efficiency story rests partly on undisclosed compute[46]. What would flip this reading: a US enforcement finding confirming a smuggled fleet above ~10,000 H100s, or a future technical report disclosing all-in training cost within ~2x of the headline. Pre-mortem: if this looks wrong in two years, the most likely reason is that the all-in estimates were too high and the efficiency even more real than credited — or, on the other side, that confirmed covert compute reframed the whole episode.
On whether an open-weight, no-moat lab is defensible: the evidence leans eroding (medium confidence). The controlling evidence is Doubao leading China at ~227M MAU to DeepSeek's ~136M[34] and documented poaching of the researchers Liang calls the moat[41], which outweighs the team-and-culture thesis because the moat's carrier — people — is exactly what better-capitalized rivals are buying. The strongest surviving counter-argument: distribution through others has worked — Azure, AWS and Nvidia hosted R1 within days, and Tencent wired it into WeChat[17][18]. What would flip this reading: DeepSeek's China model-invocation share (10.3% in H1 2025, third behind Qwen and Doubao[19]) reaching #1 in the next full-year ranking; or a closed external round that lets it counter nine-figure offers[77]. Pre-mortem: if this looks wrong in two years, the most likely reason is that ecosystem distribution substituted for owned distribution better than consumer MAU implied — or, on the other side, that a hollowed-out research bench quietly ended the release cadence.
On how much the security and IP concerns should weigh: the evidence leans documented core, inflated periphery (high confidence on the core). The controlling evidence is the verified database exposing over a million log entries[47] and Korea PIPC's finding of unlawful overseas data transfer[52], alongside a replicated 100% jailbreak rate[48] — which outweighs the "it's just geopolitics" counter because these findings are technical and regulatory, not political. The strongest surviving counter-argument: even rivals call the threat framing overstated and credit genuine independent research[44], while the China-Mobile and smuggled-chip claims rest on unconfirmed estimates[46]. What would flip this reading: the Microsoft/OpenAI distillation probe producing public evidence[42], a second verified data exposure — or, the other way, a clean re-audit from a regulator that previously sanctioned it. Pre-mortem: if this looks wrong in two years, the most likely reason is that hosted-service problems were fixed and self-hosting made them moot — or, on the other side, that a confirmed state-linked data channel proved the hawks right.
On whether it can stay at the frontier under a compute ceiling: the evidence leans fast-follower, not frontier-setter, while the ceiling holds (medium confidence). The controlling evidence is the failed full-scale Ascend training run[65] and V4 shipping ~3–6 months behind GPT-5.4 and Gemini 3.1 Pro[66], which outweighs the domestic-compute bull case because that case requires a capability — stable frontier training on Chinese silicon — that has not yet been demonstrated. The strongest surviving counter-argument: CSIS's judgment that the team would be formidable with more compute, on a track record of converting constraint into algorithmic gains[67][36]. What would flip this reading: a confirmed, fully successful Ascend or Cambricon training run for a V5-class model; or the reported first external round — a $10B+ valuation was floated in April 2026 — closing and funding a compute war chest[28]. Pre-mortem: if this looks wrong in two years, the most likely reason is that domestic silicon matured faster than 2026 reporting suggested — or, on the other side, that the gap widened past six months and "fast-follower" was itself the optimistic reading.
Sources for this section
7 sources · en, zh · 2 Chinese-language · full bibliography on the Sources list.