Open Models
Stories, products, and related signals connected to this tag in Explore.
Stories
Filter storiesOpenRouter says its free coding and sustained-agent model reached 8 trillion daily tokens within five days of launch. OpenCode separately reported processing 26 trillion Ox Alpha tokens over four days.
Community tests report Qwen 3.8 27B handling coding, OCR, and long-context workloads locally. One vLLM setup reached 3,200 tokens per minute at 262K context on two RTX 3090s without NVLink.
Local users report that Qwen 3.8 27B agent behavior changes substantially with the harness, quantization, context settings, and hardware. One 3-bit MacBook Air run at 57K context took 63 hours.
Marin has begun an open training run for a 535B-parameter mixture-of-experts model with 23B active parameters. The team plans to train on 18.75T tokens across 11 GB200 NVL72 systems over about three months.
A LocalLLaMA user reports reasoning loops caused most errors in a 2,483-task test of Qwen 3.8 27B. The report says 3–9% of inputs drove most failures because reasoning often did not terminate.
Qwen released its open-weight Qwen3.8 27B vision-language model with 262K native context and adjustable reasoning. In a 484-sample test, enabled-thinking scores fell from above 92% through 64K to 74.3–81.8% at 128K.
LocalLLaMA users said Qwen 3.8 Max 2.4T open weights are text-only while the API keeps vision support. A linked Qwen3.8-27B ModelScope page reportedly returned 404 before release.
Meta released Muse Glimmer, a 30B Apache 2.0 dense model for local agent workflows. Reports cite 4-bit builds under 20GB, vision input, function calling, 131K context, and day-0 support in Hugging Face, vLLM, SGLang, Ollama, and MLX.
Magnitude launched an open-source terminal coding agent that runs local models on-device without API keys. Its launch post says it profiles hardware and can use shell, file-editing, script, and skills tools.
Alibaba's Qwen3.8-Max is live on Venice and OpenRouter while open weights are still described as coming soon. Reports cite a 2.4T-parameter model with strong Vals and vision benchmark results.
MiniMax H3 now has Hugging Face weights, fal endpoints, AI Toolkit LoRA support, and reported single-RTX-5090 local runs. MiniMax also said deployment in the US, EU, UK, and South Korea is available through formal authorization.
Alibaba said Qwen3.8-Max left preview as a 2.4T-parameter MoE with 95B active parameters and $2/$6 per million-token pricing. Arena placed it on the Frontend Code Arena cost-performance frontier.
MiniMax released H3 weights on Hugging Face for text-to-video, image-to-video, reference-to-video, and editing workflows. vLLM-Omni, ComfyUI, SGLang Diffusion, and fal added support at launch.
DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.
Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.
MiniMax announced H3 for app and API access, with OpenRouter, Vercel, Pika, fal, Runway, and Venice integrations. Artificial Analysis ranked it No. 1 for video editing, and MiniMax said open weights are planned.
LightOn released mDenseOn and mLateOn, which translate an open English retrieval recipe into 8 languages and ship models, data, and training code. The release includes 2.8B training pairs and 16.3M fine-tuning samples.
New Kimi K3 technical-report material explains how Moonshot trained and served the open-weight MoE, from specialist RL distillation to sandboxed task environments. Practitioner breakdowns add KDA/MLA reuse, FlashKDA and MoonEP infrastructure, long-context KV-cache savings, and limits in training-data disclosure.
Anthropic opposed a categorical ban on open-weight models but backed chip controls, anti-distillation enforcement, and mandatory safety testing. The post drew backlash as more companies signed the open-weights letter.
Moonshot published Kimi K3 weights, a technical report, and a blog for a 2.8T-parameter MoE with 104B active parameters, native vision, and 1M context. The license adds separate terms for large model-as-a-service providers.
NVIDIA, Hugging Face, LangChain, Nous, Databricks, and other partners announced the Open Secure AI Alliance. The group plans to share open models, security data, agent controls, evals, and tooling.
Vercel, Sakana AI, Genspark, and Morph said they co-signed Microsoft’s Open Weights and American AI Leadership letter. The campaign drew Anthropic-focused backlash, sign-up complaints, and security-policy counterarguments.
Google, Cohere, OpenClaw, Periodic, MiniMax, and vLLM publicly backed the open-weight letter. Posts also pushed for open datasets, traces, and harnesses, while Anthropic and Amazon were noted as absent.
NVIDIA, Microsoft, Meta, Hugging Face, and other companies backed a letter defending open-weight AI. The statement argues policymakers should not conflate distillation with theft, while critics question security and openness claims.
A U.S. tech adviser accused Moonshot AI of using Anthropic’s Fable to build Kimi K3, while China’s embassy denied related claims. Engineers questioned whether leaderboard results and missing logits fit a simple distillation story.
Poolside released Laguna S 2.1, a 118B-parameter open-weight MoE with 1M context and SGLang day-zero support. Poolside and partners cite SWE-bench, Terminal-Bench, and local-agent tests.
Alibaba opened Qwen 3.8 Max Preview on Alibaba Cloud, Qwen Chat, Qoder and the web, describing it as a 2.4T model headed for open weights. Early testers praised vision results but disputed coding claims.
Posts put Kimi K3 first on Arena's Frontend Code leaderboard but 4.37-5.29 months behind U.S. frontier models in one public estimate. Other evidence cited strong DeepSWE cost-performance and cybersecurity results.
OpenBMB open-sourced MiniCPM-RobotManip, MiniCPM-RobotTrack and PhyAI, claiming local robot tracking, robot memory and throughput gains from 10 Hz to 33-36 Hz. The release packages model artifacts and a runtime path for local robot perception and manipulation experiments.
Kimi K3 posted strong coding results, including rank #5 on Artificial Analysis and #3 on DeepSWE. Engineers disputed whether its lower token price offsets higher token use and slower throughput.
Moonshot launched Kimi K3 in Kimi products and API with 1M context, native multimodality, KDA/AttnRes, and weights promised by July 27. Benchmarks place it near frontier systems, but testers cite slow serving and usability caveats.
Inkling's 1-bit GGUF ran in llama.cpp at 30–40 TPS, and TokenSpeed added day-zero support with a flat KV cache pool. Arena posts put Inkling #10 among open models in frontend code and text, while docs drew scrutiny.
Thinking Machines released Inkling with Apache 2.0 weights, 975B parameters, 41B active parameters, text/image/audio support, and up to 1M context. vLLM, SGLang, Modal, Databricks, and Vercel added day-zero support.
Testers say Kivine identifies with Moonshot/Kimi and produces strong frontend, coding, and spatial demos. Moonshot also teased Kimi K3, but the Arena claims remain unofficial.
Tencent released 1-bit and 4-bit GGUF builds for its 295B Hy3 model with llama.cpp support and MTP. Posts cite 88–92GB local runs and SWE-Bench scores of 75.4% Verified and 53.9% Pro.
A German consortium released the small SOOFI sovereign base model trained on 27T tokens. Analysts said it reuses Nemotron 3 Nano architecture and many hyperparameters with a changed data mix, and benchmarks drew criticism for overstating capability versus Qwen and Nemotron.
Tencent released Hy3 with 21B active parameters, a 256K context window, BF16/FP8 weights, and day-one vLLM/SGLang support. Kilo Code, Nous Portal, and OpenRouter also made it free for limited windows.