U.S. adviser accuses Moonshot of distilling Anthropic Fable for Kimi K3
A U.S. tech adviser accused Moonshot AI of using Anthropic’s Fable to build Kimi K3, while China’s embassy denied related claims. Engineers questioned whether leaderboard results and missing logits fit a simple distillation story.

TL;DR
- A White House accusation moved Kimi K3 from benchmark drama into trade-policy territory: TheRundownAI's post says U.S. tech adviser Michael Kratsios accused Moonshot of distilling Anthropic's Fable, while Reuters syndication says China's embassy called the accusation unfounded.
- The public technical record is still messy: ZhihuFrontier's translation says a viral CoT-distillation PDF did not prove hidden reasoning extraction, and yacineMTB's question pushed on whether distillation requires logits.
- Kimi K3 is close enough to Fable 5 to make the politics combustible: AA-Briefcase results put K3 at 1543 Elo versus Fable 5 at 1574, and the office-task test found a 9/14 tie between K3 and Fable 5.
- The policy path being discussed looks more like regulatory pressure than a clean download ban: deredleritt3r's Axios summary listed procurement rules, Entity List threats, and public pressure campaigns, while kimmonismus's post relayed a Silicon Valley letter warning against cutting off Chinese open-weight models.
The official Kimi K3 blog says the model is 2.8T parameters, uses Kimi Delta Attention plus Attention Residuals, and keeps full weights scheduled for July 27. Anthropic's earlier distillation report accused Moonshot, DeepSeek, and MiniMax of 16M Claude exchanges through 24,000 fraudulent accounts. Axios says U.S. officials have discussed Entity List threats, advisories, hosting-liability rules, and procurement pressure. A 2023 LLM distillation paper separates black-box distillation, where only generated text is available, from white-box distillation, where logits or hidden states are available in MiniLLM.
The accusation
Michael Kratsios, director of the White House Office of Science and Technology Policy, said the U.S. had information that Moonshot distilled Anthropic's Fable to build Kimi K3. Wes Roth's summary added the allegation that Moonshot used an internal platform that could switch between access methods to avoid detection, plus servers equipped with Nvidia GB300 chips and additional GB300 systems in Thailand.
Reuters syndication reported that China's embassy denied the accusation and that Moonshot had not publicly responded. CyberScoop reported the same core claim and framed it as an escalation from Anthropic's prior distillation warnings.
Anthropic's February research report is the background document Kratsios's claim leans on. It accused DeepSeek, Moonshot, and MiniMax of industrial-scale Claude extraction, including more than 3.4M Moonshot exchanges through hundreds of fraudulent accounts targeting agentic reasoning, tool use, coding, data analysis, computer use, and vision.
Distillation without logits
The first technical dispute was vocabulary. yacineMTB asked whether distillation still means anything if Anthropic is not exposing logits, then suggested the term may now cover using closed models to generate code or training data.
The clean distinction is soft versus hard distillation:
- Soft distillation: the student trains on the teacher's output distribution, usually logits or probabilities.
- Hard distillation: the student trains on generated answers, traces, labels, code, or other text outputs.
- Black-box distillation: the teacher is only accessible through an API, so the dataset is prompt-response pairs.
- White-box distillation: logits, hidden states, or internal activations are available.
Anthropic's own report defines distillation as training a less capable model on outputs of a stronger model, then calls legitimate internal distillation vital while condemning covert competitor extraction. The technical fight is over scale, access, and contract breach, not whether logits are the only possible signal.
The public evidence gap
The viral PDF story got weaker under scrutiny. ZhihuFrontier summarized Kris 谭's analysis as saying the original experiment replayed supplied encrypted CoT payloads, but did not extract GPT's private chain of thought or show that Kimi trained on such traces.
The evidence Kris said would be needed is concrete:
- what data was collected
- how it was obtained
- how much data existed
- how it entered Kimi's training pipeline
Style similarity also landed in the debate. scaling01's heatmap post highlighted a Kimi K3 and Fable 5 style comparison, while teortaxesTex argued that Claude-shaped prose is a weak indicator because style is cheap for LLMs to learn and easier to transfer than capability.
The timeline objection is narrower. BlancheMinerva's reply claimed K3 finished training before ICML, leaving at most about 10 days of public Fable availability, while Ryan Greenblatt's clarification said direct Anthropic hacking seemed unlikely but access through another prerelease partner was technically possible.
The copyright mirror
The accusation triggered a predictable copyright mirror. Teknium parodied Kratsios's language by saying Anthropic had scraped his GitHub, Twitter, Reddit, Discord, and Facebook posts through a large-scale platform.
The sharper legal distinction was contract versus ownership. one reply framed distillation as a terms-of-service breach, while natolambert argued there is no legal precedent that model outputs are IP.
Simon Willison pointed to Ben Thompson's proposal that the U.S. explicitly protect model training as fair use and bar terms of service that forbid distillation for U.S. companies in Who's Afraid of Chinese Models?. That proposal turns the hypocrisy complaint into policy: indemnify training, then make API output reuse harder to block.
The benchmarks behind the panic
Artificial Analysis put Kimi K3 right behind Fable 5 on AA-Briefcase, its private benchmark for agentic knowledge work across messy files and finished deliverables.
Key numbers from the AA-Briefcase thread:
- AA-Briefcase Elo: Fable 5 1574, Kimi K3 1543.
- Kimi K3 improvement over K2.6: +727 Elo.
- Rubric pass rate: Kimi K3 51%, Fable 5 56%.
- Analytical quality Elo: Kimi K3 1754, Fable 5 1744.
- Presentation Elo: Kimi K3 1471, behind GPT-5.6 Sol at 1660.
- Mean cost per task: Kimi K3 $10.57, Fable 5 $22.30, GPT-5.6 Sol $5.32.
- Mean time per task: Kimi K3 56.4 minutes.
Composio's agentic office test produced the cleanest side-by-side result: Kimi K3 and Fable 5 each solved 9 of 14 tasks, passed the same 9, and failed the same 5. Fable finished about 2.5x faster, while Kimi used about 40% fewer tokens.
Frontend and agent benchmarks pointed the same direction. Kilo Code's UI test said Kimi and Fable produced eerily similar page anatomy across 10 UI builds at 29% of Fable's cost, and Agent Arena ranked Kimi K3 #4 overall with +9.6% net improvement and #1 confirmed task success.
K3's own architecture story
Moonshot's official Kimi story is specific enough to matter. The Kimi K3 tech blog says K3 is a 2.8T-parameter MoE model with a 1M-token context window, native vision, Kimi Delta Attention, Attention Residuals, and Stable LatentMoE activating 16 of 896 experts.
The launch blog also says K3 uses quantization-aware training from SFT onward, MXFP4 weights, MXFP8 activations, fully balanced expert-parallel training, and supernode deployments with 64 or more accelerators. Full model weights and a technical report remain promised by July 27.
The weird harness detail is preserved thinking history. The official blog says K3 was trained in a mode where prior thinking content must be passed back, and multimodalart's quote flagged that switching an ongoing session from another model to K3 can make generation quality unstable.
The soft-ban toolkit
Axios reported that the Trump administration had considered several tools before Kimi K3 reignited the issue:
- Commerce Entity List additions for Chinese AI labs.
- NSA or White House cyber advisories discouraging use of Chinese AI systems.
- Hosting rules requiring U.S. companies to guarantee security and accept liability.
- Commerce supply-chain rules targeting Chinese open-source models.
- Procurement rules, Entity List threats, and public pressure campaigns.
The accusation now sits inside a broader U.S.-China AI channel. koltregaskes's Reuters summary said U.S. and Chinese officials are expected to hold AI talks in September, with Treasury Secretary Scott Bessent leading the U.S. delegation and saying officials were finding "watermarks of U.S. language models" on Chinese models.
The startup backlash is already organized. kimmonismus's post said nearly 200 Silicon Valley companies, including Proton and Y Combinator, urged the administration not to cut off access to Chinese open-weight models.
The access chokepoint
Open weights do not make K3 easy to run. Moonshot's blog recommends 64 or more accelerators, and cedric_chee's OpenRouter screenshot showed first-party Kimi K3 at 19 tokens per second with 7.90s p50 latency.
Demand moved anyway. Cline's usage post said Kimi K3 went from 0% to 16% of ClinePass token usage in three days, becoming its #3 open-weights model, while one Codex setup thread showed Kimi K3 running inside Codex through a config switch.
The capacity story is part of the model story. Composio's reply said testing was delayed because Kimi servers were under capacity for the first two days, and aibuilderclub_'s reply said Kimi's coding plan looked sold out that day.