Skip to content
AI Primer
MODEL21 stories

Kimi

Stories, products, and related signals connected to this tag in Explore.

NEWS8th August
Kimi K3 reportedly reaches GitHub after benchmark sandbox leaves outbound access open

Frontier Security reportedly ran public Kimi K3 in an open-source cyber sandbox and saw it reach GitHub after outbound network access was left open. The UK AI Security Institute said it did not run the test.

NEWS8th August
Together AI ranks first or tied first on 3 of 4 Kimi K3 provider benchmarks

Together AI said it ranked first or tied first on three of four Kimi K3 provider benchmarks, while Baseten described a 2.8T-parameter Blackwell GB300 serving stack. Local users also reported trimming the model from 711GB to 478GB and running it through llama.cpp RPC across clusters.

RELEASE1w ago
Wafer launches Kimi K3 Fast on OpenRouter with 172 output tokens/sec claim

Wafer listed Kimi K3 Fast on OpenRouter and Vercel AI Gateway. It claimed 172 output tokens/sec, 15.8s end-to-end latency, and provider routing through OpenRouter’s :nitro option.

NEWS2w ago
Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing

Composio found Kimi K3 success stayed near 71–79% across three harnesses while median token use ranged from 61K to 340K and cost from $0.22 to $2 per task. Cline also shipped a Kimi K3 CLI update.

NEWS2w ago
Kimi K3 report details RL distillation and FlashKDA infrastructure

New Kimi K3 technical-report material explains how Moonshot trained and served the open-weight MoE, from specialist RL distillation to sandboxed task environments. Practitioner breakdowns add KDA/MLA reuse, FlashKDA and MoonEP infrastructure, long-context KV-cache savings, and limits in training-data disclosure.

RELEASE2w ago
Kimi K3 launches across vLLM, SGLang, Ollama, and OpenRouter

Kimi K3 landed in major serving stacks on launch day, including vLLM, SGLang, Ollama, OpenRouter, Fireworks, Together, Modal, and Vercel AI Gateway. Providers cited ZDR options, optimization work, and prices around $3/M input and $15/M output.

RELEASE2w ago
Moonshot releases Kimi K3 open weights with 2.8T-parameter MoE

Moonshot published Kimi K3 weights, a technical report, and a blog for a 2.8T-parameter MoE with 104B active parameters, native vision, and 1M context. The license adds separate terms for large model-as-a-service providers.

NEWS3w ago
U.S. adviser accuses Moonshot of distilling Anthropic Fable for Kimi K3

A U.S. tech adviser accused Moonshot AI of using Anthropic’s Fable to build Kimi K3, while China’s embassy denied related claims. Engineers questioned whether leaderboard results and missing logits fit a simple distillation story.

NEWS3w ago
Artificial Analysis reports Kimi K3 averages 56.4 minutes on AA-Briefcase

Artificial Analysis reports Kimi K3 averages 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task. Kilo also found UI-build outputs close to Claude Fable 5 at 29% of the cost.

NEWS3w ago
Moonshot pauses new Kimi K3 subscriptions after GPU capacity crunch

Moonshot said Kimi K3 demand pushed its GPUs near capacity, so it paused new subscriptions and split memberships into Kimi and Kimi Code plans. Users also reported slow serving and sold-out paid plans.

NEWS3w ago
Kimi K3 ranks No. 1 on Arena Frontend Code leaderboard

Posts put Kimi K3 first on Arena's Frontend Code leaderboard but 4.37-5.29 months behind U.S. frontier models in one public estimate. Other evidence cited strong DeepSWE cost-performance and cybersecurity results.

NEWS3w ago
Kimi K3 benchmarks last at 53/67 in AlphaSignal repair harness

AlphaSignal's repair harness put Kimi K3 last at 53/67, while other tests ranked it high on DeepSWE, Vibe Code Bench, legal work, cyber, and CUDA kernels. Cost often beat Fable, but speed lagged.

NEWS3w ago
Red-teamers claim Kimi K3 jailbreaks produced cyber and bio outputs

Multiple posts claimed Kimi K3 jailbreaks produced harmful cyber and bio-related outputs. Other users asked for setups or pointed to UK and US cyber ranges as better tests of real capability.

NEWS3w ago
Kimi K3 ranks #5 on Artificial Analysis as engineers dispute coding cost

Kimi K3 posted strong coding results, including rank #5 on Artificial Analysis and #3 on DeepSWE. Engineers disputed whether its lower token price offsets higher token use and slower throughput.

RELEASE3w ago
Moonshot launches Kimi K3 with 2.8T parameters and 1M context

Moonshot launched Kimi K3 in Kimi products and API with 1M context, native multimodality, KDA/AttnRes, and weights promised by July 27. Benchmarks place it near frontier systems, but testers cite slow serving and usability caveats.

NEWS4w ago
Users report Kivine on LMArena may be a Kimi K3 preview

Testers say Kivine identifies with Moonshot/Kimi and produces strong frontend, coding, and spatial demos. Moonshot also teased Kimi K3, but the Arena claims remain unofficial.

RELEASE1mo ago
Cognition launches SWE-1.7 in Devin at 1,000 tok/s

Cognition says SWE-1.7 was trained with RL on a Kimi K2.7 base and now runs in Devin at 1,000 tok/s. It reports 42.3% on FrontierCode at $1.97 per task and released revised grading rules.

RELEASE2mo ago
Moonshot releases Kimi K2.7 Code: +21.8% on Kimi Code Bench v2, 30% fewer reasoning tokens

Moonshot open-sourced Kimi K2.7 Code and says it outperforms K2.6 by 21.8% on Kimi Code Bench v2 while using 30% fewer reasoning tokens. The release includes open weights and API access, so teams can test the 180 tok/s HighSpeed rollout and early Cline/OpenCode support.

NEWS3mo ago
Kimi K2.6 adds free Hermes and Cline access plus Replicate, Perplexity, and Together support

A day after Kimi K2.6’s launch, providers and tools opened new access paths including temporary free use in Hermes and Cline plus availability on Replicate, Together, Perplexity, and Tinker. Engineers can test the open model across agent harnesses and hosted runtimes without standing up their own stack first.

RELEASE3mo ago
Kimi K2.6 launches with 58.6 SWE-Bench Pro and 4,000-tool-call agent runs

Moonshot open-sourced Kimi K2.6, a 1T-parameter MoE with 32B active parameters, 256K context, multimodal input, and larger agent swarms. It now sits near frontier closed models for long-horizon coding and tool use, so teams can try it for agent workflows.

NEWS3mo ago
Kimi K2.6 adds day-one support across vLLM, SGLang, Ollama, and OpenRouter

Kimi K2.6 shipped across vLLM, SGLang, OpenRouter, Baseten, Ollama, OpenCode, Hermes Agent, and Droid within hours of launch. That cuts the usual lag between model release and production trials, so mixed-provider agent stacks can test it sooner.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.