Open Models
Stories, products, and related signals connected to this tag in Explore.
Stories
Filter storiesPerplexity's open pplx-embed-v2-late models retrieve text and images in a shared multi-vector space. The 0.6B model can query indexes built by the 9B model, including document pages without OCR.
Beam is a sparse mixture-of-experts model with 23B active parameters, a reported 1M-token effective context, and Apache 2.0 weights. Independent benchmarking is beginning.
Aleph Alpha released Kolibri, an Apache 2.0 open-weight MoE with a 1M-token context window. The model has 78B total parameters and 3.46B active parameters, plus reasoning, tool calling and German-focused training.
Kev 1.0 introduces Kev-27B and updates Kev-9B with a 64k document window and TypeSafe SDK compatibility. The release includes weights, source code, and fine-tuning and deployment skills.
Perplexity released pplx-embed-v2-context-9b-preview, which encodes chunks using whole-document context. Perplexity reports leading results on ConTEB and Turbopuffer's context benchmark.
Holo4 ships as a dense 27B model and a 35B-A3B mixture-of-experts model, with API access and downloadable weights. H Company reports a 61.7% score on OSWorld 2.0 and publishes replayable benchmark trajectories.
Experiments reported by Maxime Rivest found task-specific small classifiers outperforming frontier models on specialized decisions. One result says a 4-million-parameter BERT Tiny model beat Opus and Kimi after training on 10,000 examples.
CUA released Cua-S1-4B-0.2, a multimodal decision model trained with supervised learning and task-completion RL in live environments. CUA reports 92.9% on a frozen GUI-360 split and released adapters and training code.
Black Forest Labs released FLUX 3 Action, an open 7B model that jointly predicts future video and actions for robot policies. The company reports first place on RoboLab and released embodiment fine-tunes, training recipes, and Jetson deployment support.
Xiaomi released MiMo V2.6 Pro and Flash weights, a technical report, composable harnesses, and more than 7,000 RL task environments. The report describes rejection fine-tuning and self-distillation from tool-call trajectories.
Parakeet Redux compresses NVIDIA's Parakeet from 1.2 GB to 178 MB with ternary weights. Its author reports 113× real-time CPU speed and stronger results on the 25-language FLEURS benchmark.
Xiaomi released Pro and Flash MiMo-V2.6 mixture-of-experts models with open weights and a 1M-token context window. The release includes an RL dashboard and day-one vLLM support, while training artifacts are planned.
Qwen-Image-2.1 is a 7B model that combines image generation and editing, including native RGBA output and support for up to 10 reference images. It has day-one support in ComfyUI, Diffusers, vLLM-Omni, and Ostris, but its license is non-com
Kev is an Apache-2.0 family of 0.6B, 4B, and 8B decision models compatible with TypeSafe System One APIs. Its author reports that the 8B model reached 79.6% out-of-domain accuracy versus Jev’s 85.7%, while the 4B model runs on a 32 GB Mac.
TypeSafe says jev-use generates candidate actions from browser state, has Jev select one, then validates and executes it through Cua Driver. Its CUA-S1-FORMS model scored forms locally in 7–9 ms, excluding execution.
CUA open-sourced CUA-S1-FORMS, a specialist model that selects bounded actions such as filling fields, checking boxes, clicking, or skipping. Cua Driver executes the ordered plan, and the MIT release includes synthetic-data generation, training, evaluation, and deployment tools.
DeepSeek V4.1 Flash leads Vals and Artificial Analysis open-weight comparisons, according to the evaluators. Its encoder-decoder design shares compressed KV state across decoder layers to reduce serving costs.
DeepSeek released V4.1 Flash, a 552B-parameter multimodal MoE model with 8B input and 16B output active parameters. It supports up to 1 million tokens of context, while an independent BridgeBench run used 23.5 million tokens on one task.
Cohere released an open-source serving system that fuses the LLM decode step into one GPU kernel launch. On North Mini Code with one H100, it reports up to 1.58x vLLM performance at the tested batch size.
Magic says a new pretraining recipe matched DeepSeek V4 Pro with roughly 50 times less compute. After a 10x scale-up costing about $4 million, the company says it exceeded publicly available base models.
OpenBMB released MiniCPM5-2B under Apache 2.0 with its data, training recipes, and RL stack. The model has day-zero deployment support in vLLM and SGLang.
MBZUAI released K2 Horizon models from 0.9B to 375B parameters with weights, training code, data recipes, intermediate checkpoints, logs, and evaluations. vLLM says the family has day-zero support and up to 512K context.
Open Athena has begun training Marin, a 535B-parameter MoE with 23B active parameters, over 18T tokens. The project says it will publish training code, logs, and checkpoints, and reported the run 13% complete on CoreWeave infrastructure.
Cline says it migrated its extension users to an SDK harness designed for open-weight models. It reports task mistake rates fell from 6.34% to 0.62%.
Together AI says its HUMAIN partnership will provide 250 MW of data-center capacity for open-source models over the next year. The deployment exceeds 100,000 chips and could produce more than $5 billion in annualized AI.
A new account says the incident involved multiple waves of agents. Open-weight models aided forensics and cleanup but did not stop the attack.
Tencent released Hy4 Preview, a 770B-parameter mixture-of-experts model with 49B active parameters and a 1M-token context window. vLLM added day-zero support, while Cline, OpenCode Go, and Vercel AI Gateway made the model available.
Z.ai released GLM-5.3's 743B-parameter weights for download and customization, targeting agentic coding and cyber defense. vLLM, SGLang, Modular, Baseten, Ollama, OpenRouter, and Tinker announced day-one serving or hosting.
Qwen released Qwen3.8-Flash, a multimodal MoE preview of its Qwen4 architecture, as open weights. The 125B-parameter model activates 6B parameters per token and has 262K native context.
Perceptron released weights, inference code, and training details for Isaac 0.5, an embodied model for video perception, reasoning, and robot control. The 36B dynamic-MoE model was trained on 1 million hours of video.
Z.ai identified the formerly anonymous Ox Alpha as GLM-5.3-Flash and released it under an MIT license. The native multimodal model has 320B total parameters, 18B active parameters, and a 1M-token context window.
Perplexity’s Portable Computer runs its orchestrator, subagents, and harness locally on NVIDIA DGX Spark with a post-trained 27B model. Frontier-model escalation requires user approval and flags PII before text is sent externally.
OpenRouter says its free coding and sustained-agent model reached 8 trillion daily tokens within five days of launch. OpenCode separately reported processing 26 trillion Ox Alpha tokens over four days.
Community tests report Qwen 3.8 27B handling coding, OCR, and long-context workloads locally. One vLLM setup reached 3,200 tokens per minute at 262K context on two RTX 3090s without NVLink.
Marin has begun an open training run for a 535B-parameter mixture-of-experts model with 23B active parameters. The team plans to train on 18.75T tokens across 11 GB200 NVL72 systems over about three months.
Local users report that Qwen 3.8 27B agent behavior changes substantially with the harness, quantization, context settings, and hardware. One 3-bit MacBook Air run at 57K context took 63 hours.
A LocalLLaMA user reports reasoning loops caused most errors in a 2,483-task test of Qwen 3.8 27B. The report says 3–9% of inputs drove most failures because reasoning often did not terminate.
Qwen released its open-weight Qwen3.8 27B vision-language model with 262K native context and adjustable reasoning. In a 484-sample test, enabled-thinking scores fell from above 92% through 64K to 74.3–81.8% at 128K.
LocalLLaMA users said Qwen 3.8 Max 2.4T open weights are text-only while the API keeps vision support. A linked Qwen3.8-27B ModelScope page reportedly returned 404 before release.
Meta released Muse Glimmer, a 30B Apache 2.0 dense model for local agent workflows. Reports cite 4-bit builds under 20GB, vision input, function calling, 131K context, and day-0 support in Hugging Face, vLLM, SGLang, Ollama, and MLX.
Magnitude launched an open-source terminal coding agent that runs local models on-device without API keys. Its launch post says it profiles hardware and can use shell, file-editing, script, and skills tools.
Alibaba's Qwen3.8-Max is live on Venice and OpenRouter while open weights are still described as coming soon. Reports cite a 2.4T-parameter model with strong Vals and vision benchmark results.
MiniMax H3 now has Hugging Face weights, fal endpoints, AI Toolkit LoRA support, and reported single-RTX-5090 local runs. MiniMax also said deployment in the US, EU, UK, and South Korea is available through formal authorization.
MiniMax released H3 weights on Hugging Face for text-to-video, image-to-video, reference-to-video, and editing workflows. vLLM-Omni, ComfyUI, SGLang Diffusion, and fal added support at launch.
Alibaba said Qwen3.8-Max left preview as a 2.4T-parameter MoE with 95B active parameters and $2/$6 per million-token pricing. Arena placed it on the Frontend Code Arena cost-performance frontier.
DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.
Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.
LightOn released mDenseOn and mLateOn, which translate an open English retrieval recipe into 8 languages and ship models, data, and training code. The release includes 2.8B training pairs and 16.3M fine-tuning samples.
MiniMax announced H3 for app and API access, with OpenRouter, Vercel, Pika, fal, Runway, and Venice integrations. Artificial Analysis ranked it No. 1 for video editing, and MiniMax said open weights are planned.
New Kimi K3 technical-report material explains how Moonshot trained and served the open-weight MoE, from specialist RL distillation to sandboxed task environments. Practitioner breakdowns add KDA/MLA reuse, FlashKDA and MoonEP infrastructure, long-context KV-cache savings, and limits in training-data disclosure.