Liang Wenfeng reportedly frames DeepSeek roadmap around scarce GPU supply
Posts quoting Liang Wenfeng said DeepSeek is targeting low positive API margins while constrained by GPU supply. They also said early-June capacity was about 20,000 H100-equivalent units and that Huawei capacity remains below Nvidia.

TL;DR
- GPU supply is the roadmap constraint: teortaxesTex quoted Liang Wenfeng saying DeepSeek had about 20,000 H100-equivalent units in early June and wanted to spend everything available within six months, “basically all NVIDIA.”
- DeepSeek’s model is open weights plus thin API margins: teortaxesTex said Wenfeng defined “reasonable profit” as recouping compute acquisition cost within 10 months, while teortaxesTex's quote roundup framed the commitment as always open sourcing the best model.
- Huawei capacity is useful but not enough for the next scale step: teortaxesTex quoted Wenfeng saying Huawei could provide about 16,000 GPUs, and teortaxesTex's capacity quote said that equals only about 4,000 B-series GPUs.
- V4 still looks like a transition model: teortaxesTex said the updated V4 should have native multimodality, while teortaxesTex's delay note said V4 GA slipped from a hoped-for June target and a 150B-active training run was pushed toward late 2026 to April 2027.
- The funding story is partly a retention story: teortaxesTex argued DeepSeek’s compute spend through 2026 could be covered by Liang’s personal investment, with the financing round mostly functioning as equity to keep people from leaving.
DeepSeek’s own V4 release note gives the backdrop: V4-Pro is listed at 1.6T total parameters with 49B active, V4-Flash at 284B total with 13B active, both with 1M context and open weights. The official pricing page puts V4-Pro at $0.003625 per 1M cache-hit input tokens, $0.435 cache-miss input, and $0.87 output. The TileKernels repo makes Wenfeng’s CUDA-moat comments less hand-wavy: DeepSeek says some TileLang kernels are already used in internal training and inference.
Open weights, thin margins
The investor-call transcript, as summarized by teortaxesTex, puts open source at the center of DeepSeek’s strategy. Wenfeng reportedly told investors that open sourcing is “the agenda,” and that those not aligned could buy more Zhipu instead.
The stronger claim is narrower than the meme. teortaxesTex's quote has Wenfeng saying DeepSeek’s strongest models will “probably” be open sourced, while teortaxesTex's follow-up reads the commitment as not deploying a stronger first-party model than the one released.
The margin target is just as concrete: recoup compute acquisition cost in 10 months, then avoid both subsidies and excess rent. DeepSeek’s pricing page is the public surface of that strategy, with V4-Flash output at $0.28 per 1M tokens and V4-Pro output at $0.87.
That is why engineers keep benchmarking DeepSeek economics against everything else. Few frontier labs publish enough operational detail for the comparison to be this annoying.
No Doubao superapp
Wenfeng reportedly told investors he could have pursued a Doubao-style superapp, but a ByteDance fight was not worth it when AGI was the larger prize. The product posture is API-first and infrastructure-first, not consumer-app maximalist.
teortaxesTex’s earlier framing was sharper: Wenfeng treats intelligence as an industrial input, closer to electricity than software rent. teortaxesTex's 2024 quote attributed the logic to Liang’s view that the endgame is higher social efficiency and a company’s place in the industrial division chain.
The China angle follows from the same premise. teortaxesTex's production quote had Wenfeng saying Chinese companies are likely to be the largest producer in the global AI division of labor, citing chip capacity, electricity, and manufacturing capacity.
The 20K H-equivalent ceiling
DeepSeek’s reported early-June capacity was about 20,000 H100-equivalent units. Wenfeng reportedly wanted to turn every available cent into GPUs within six months, and most of that near-term spend meant NVIDIA.
The bigger constraint was availability, not desire. teortaxesTex's compute note said Wenfeng did not expect to spend more than 20 billion yuan on compute in 2026 and gave a hardware translation for very large training runs: roughly 50,000 GB300s or 200,000 Huawei 950s for a model at about 800B activated parameters.
The near-term roadmap is smaller. teortaxesTex's activation-scale quote has Wenfeng saying DeepSeek would not compete with the U.S. at that scale yet, and would move toward 150B, 156B, or 250B activated-parameter models when resources improved.
Huawei 950s and the CUDA moat
Wenfeng’s domestic-compute claim has two parts that should not be collapsed:
- Ecosystem: teortaxesTex's domestic-compute excerpt quoted him saying domestic AI chips have no hardware or ecosystem problem, only insufficient production capacity.
- CUDA replacement: the same excerpt listed three reasons NVIDIA’s CUDA moat is being dismantled: AI can write code, DeepSeek has TileLang-style tooling, and NVIDIA’s GPU lineage carries gaming-era design baggage.
- Supernode parity: teortaxesTex's Huawei comparison quoted him saying a Huawei 950 supernode can substitute for GB200 and GB300 on performance and price.
- Chip-level gap: that same quote still put the gap at four Huawei GPUs per NVIDIA GPU, plus about two years.
DeepSeek’s TileKernels repo backs the software side of the claim: it describes TileLang as a Python DSL for high-performance GPU kernels and says most kernels approach hardware limits for compute intensity and memory bandwidth.
Capacity is the catch. teortaxesTex on Huawei capacity quoted Wenfeng saying Huawei could provide DeepSeek capacity for about 16,000 GPUs, possibly all Huawei had available, and teortaxesTex's 16,000-GPU quote said that amount was enough for the current generation but not enough for a next-generation model.
Funding as retention and compute prepayment
36Kr's funding note reported DeepSeek’s first round at about 51 billion yuan, with Liang contributing about 20 billion yuan as the largest single investor. That lines up with teortaxesTex’s read that Liang’s personal investment could cover compute through 2026.
The same post said DeepSeek was not buying data and framed the round as equity to prevent departures. 36Kr's infrastructure story also reported new data-center hiring, including IDC planning roles, after the financing.
The strangest reported line was profitability through scarcity. teortaxesTex's net-profit note quoted the idea that DeepSeek might already be at net profit because the GPUs it wants to buy do not exist in sufficient supply.
V4 update path
V4 was reportedly as large as DeepSeek could train at the time. The updated version should have native multimodality, according to teortaxesTex, while V4 GA reportedly missed a June target.
The DeepSeek V4 paper describes the preview family as two MoE models with 1M context, hybrid attention using CSA and HCA, Manifold-Constrained Hyper-Connections, Muon optimization, and more than 32T pretraining tokens. The DSpark paper shows where the “bang for the buck” work is going now: speculative decoding for DeepSeek-V4 serving under live traffic, focused on reducing verification waste in high-concurrency systems.
The next generation is supposed to add continual learning. Before that, teortaxesTex expects more cost-performance iteration, and teortaxesTex's cost quote quoted Wenfeng saying lower algorithmic cost lets DeepSeek train and afford bigger models.
The AGI sequence is explicit but staged. teortaxesTex's roadmap quote summarized it as learning to learn, then a self-iterating intelligence singularity, then embodied intelligence; teortaxesTex's world-model note said Wenfeng was skeptical of world models.
Data annotation and self-use
DeepSeek’s labor bottleneck may be less glamorous than its GPU bottleneck. teortaxesTex's annotation note said China had no cost advantage in data annotation and that half of DeepSeek, including core researchers, was spending time on it.
The timing is recent: teortaxesTex's follow-up quoted Wenfeng saying China had started high-quality data annotation only in the last half-year. teortaxesTex tied that to reduced jaggedness in GLM 5.2 and K3, with distillation covering the rest.
DeepSeek’s internal target is also unusually self-referential. teortaxesTex's self-use quote said the first goal of DeepSeek’s models is not that users find them good, but that DeepSeek finds them good to use internally.
Hallucinations sit outside the main push. teortaxesTex's hallucination quote said Wenfeng classifies hallucinations as a product problem, solvable long term but not the current focus, while teortaxesTex's capability note said he had no plan to beat American AI on raw capability apart from cost.