Skip to content
AI Primer
TOPIC50 stories

Open Models

Stories, products, and related signals connected to this tag in Explore.

RELEASE7th October
Perplexity releases pplx-embed-v2-late models for shared text-image retrieval

Perplexity's open pplx-embed-v2-late models retrieve text and images in a shared multi-vector space. The 0.6B model can query indexes built by the 9B model, including document pages without OCR.

RELEASE5th October
Reflection releases 501B-parameter Beam model

Beam is a sparse mixture-of-experts model with 23B active parameters, a reported 1M-token effective context, and Apache 2.0 weights. Independent benchmarking is beginning.

RELEASE3rd October
Aleph Alpha releases Kolibri with a 1M-token context window

Aleph Alpha released Kolibri, an Apache 2.0 open-weight MoE with a 1M-token context window. The model has 78B total parameters and 3.46B active parameters, plus reasoning, tool calling and German-focused training.

RELEASE1w ago
Kev releases 1.0 open-weight decision models with a 64k document window

Kev 1.0 introduces Kev-27B and updates Kev-9B with a 64k document window and TypeSafe SDK compatibility. The release includes weights, source code, and fine-tuning and deployment skills.

RELEASE1w ago
Perplexity open-sources pplx-embed-v2-context-9b-preview

Perplexity released pplx-embed-v2-context-9b-preview, which encodes chunks using whole-document context. Perplexity reports leading results on ConTEB and Turbopuffer's context benchmark.

RELEASE1w ago
H Company releases 27B and 35B-A3B Holo4 computer-use models

Holo4 ships as a dense 27B model and a 35B-A3B mixture-of-experts model, with API access and downloadable weights. H Company reports a 61.7% score on OSWorld 2.0 and publishes replayable benchmark trajectories.

WORKFLOW2w ago
4M-parameter BERT Tiny reportedly beats Opus and Kimi after training on 10,000 examples

Experiments reported by Maxime Rivest found task-specific small classifiers outperforming frontier models on specialized decisions. One result says a 4-million-parameter BERT Tiny model beat Opus and Kimi after training on 10,000 examples.

RELEASE2w ago
CUA releases Cua-S1-4B-0.2 for computer use

CUA released Cua-S1-4B-0.2, a multimodal decision model trained with supervised learning and task-completion RL in live environments. CUA reports 92.9% on a frozen GUI-360 split and released adapters and training code.

RELEASE2w ago
Black Forest Labs releases open 7B FLUX 3 Action model

Black Forest Labs released FLUX 3 Action, an open 7B model that jointly predicts future video and actions for robot policies. The company reports first place on RoboLab and released embodiment fine-tunes, training recipes, and Jetson deployment support.

RELEASE2w ago
Xiaomi releases MiMo V2.6 Pro and Flash model weights

Xiaomi released MiMo V2.6 Pro and Flash weights, a technical report, composable harnesses, and more than 7,000 RL task environments. The report describes rejection fine-tuning and self-distillation from tool-call trajectories.

RELEASE2w ago
Parakeet Redux cuts NVIDIA's Parakeet from 1.2 GB to 178 MB

Parakeet Redux compresses NVIDIA's Parakeet from 1.2 GB to 178 MB with ternary weights. Its author reports 113× real-time CPU speed and stronger results on the 25-language FLEURS benchmark.

RELEASE2w ago
Xiaomi releases open-weight MiMo-V2.6 models with 1M-token context

Xiaomi released Pro and Flash MiMo-V2.6 mixture-of-experts models with open weights and a 1M-token context window. The release includes an RL dashboard and day-one vLLM support, while training artifacts are planned.

RELEASE2w ago
Qwen releases Qwen-Image-2.1 with open weights

Qwen-Image-2.1 is a 7B model that combines image generation and editing, including native RGBA output and support for up to 10 reference images. It has day-one support in ComfyUI, Diffusers, vLLM-Omni, and Ostris, but its license is non-com

RELEASE2w ago
Kev releases open decision models built on Qwen3

Kev is an Apache-2.0 family of 0.6B, 4B, and 8B decision models compatible with TypeSafe System One APIs. Its author reports that the 8B model reached 79.6% out-of-domain accuracy versus Jev’s 85.7%, while the 4B model runs on a 32 GB Mac.

RELEASE3w ago
TypeSafe integrates Jev with Cua Driver for browser actions

TypeSafe says jev-use generates candidate actions from browser state, has Jev select one, then validates and executes it through Cua Driver. Its CUA-S1-FORMS model scored forms locally in 7–9 ms, excluding execution.

RELEASE3w ago
CUA releases open-source CUA-S1-FORMS for bounded web-form actions

CUA open-sourced CUA-S1-FORMS, a specialist model that selects bounded actions such as filling fields, checking boxes, clicking, or skipping. Cua Driver executes the ordered plan, and the MIT release includes synthetic-data generation, training, evaluation, and deployment tools.

RELEASE4w ago
DeepSeek V4.1 Flash tops independent open-weight evaluations

DeepSeek V4.1 Flash leads Vals and Artificial Analysis open-weight comparisons, according to the evaluators. Its encoder-decoder design shares compressed KV state across decoder layers to reduce serving costs.

RELEASE4w ago
DeepSeek releases 552B-parameter V4.1 Flash multimodal model

DeepSeek released V4.1 Flash, a 552B-parameter multimodal MoE model with 8B input and 16B output active parameters. It supports up to 1 million tokens of context, while an independent BridgeBench run used 23.5 million tokens on one task.

RELEASE4w ago
Cohere open-sources fused LLM decode kernel with 1.58x vLLM claim

Cohere released an open-source serving system that fuses the LLM decode step into one GPU kernel launch. On North Mini Code with one H100, it reports up to 1.58x vLLM performance at the tested batch size.

NEWS4w ago
Magic says its pretraining recipe matches DeepSeek V4 Pro with 50x less compute

Magic says a new pretraining recipe matched DeepSeek V4 Pro with roughly 50 times less compute. After a 10x scale-up costing about $4 million, the company says it exceeded publicly available base models.

RELEASE4w ago
OpenBMB releases 2.6B-parameter MiniCPM5-2B under Apache 2.0

OpenBMB released MiniCPM5-2B under Apache 2.0 with its data, training recipes, and RL stack. The model has day-zero deployment support in vLLM and SGLang.

RELEASE1mo ago
MBZUAI releases six K2 Horizon models from 0.9B to 375B parameters

MBZUAI released K2 Horizon models from 0.9B to 375B parameters with weights, training code, data recipes, intermediate checkpoints, logs, and evaluations. vLLM says the family has day-zero support and up to 512K context.

NEWS1mo ago
Open Athena begins training 535B-parameter Marin model

Open Athena has begun training Marin, a 535B-parameter MoE with 23B active parameters, over 18T tokens. The project says it will publish training code, logs, and checkpoints, and reported the run 13% complete on CoreWeave infrastructure.

RELEASE1mo ago
Cline migrates 11 million extension users to an SDK harness

Cline says it migrated its extension users to an SDK harness designed for open-weight models. It reports task mistake rates fell from 6.34% to 0.62%.

NEWS1mo ago
Together AI signs 250 MW HUMAIN capacity deal for open models

Together AI says its HUMAIN partnership will provide 250 MW of data-center capacity for open-source models over the next year. The deployment exceeds 100,000 chips and could produce more than $5 billion in annualized AI.

NEWS1mo ago
Hugging Face says it contained agent backdoors after a days-long response

A new account says the incident involved multiple waves of agents. Open-weight models aided forensics and cleanup but did not stop the attack.

RELEASE1mo ago
Tencent releases 770B-parameter Hy4 Preview open weights

Tencent released Hy4 Preview, a 770B-parameter mixture-of-experts model with 49B active parameters and a 1M-token context window. vLLM added day-zero support, while Cline, OpenCode Go, and Vercel AI Gateway made the model available.

RELEASE1mo ago
Z.ai releases 743B-parameter GLM-5.3 open weights

Z.ai released GLM-5.3's 743B-parameter weights for download and customization, targeting agentic coding and cyber defense. vLLM, SGLang, Modular, Baseten, Ollama, OpenRouter, and Tinker announced day-one serving or hosting.

RELEASE1mo ago
Qwen releases Qwen3.8-Flash open weights: 125B MoE with 262K context

Qwen released Qwen3.8-Flash, a multimodal MoE preview of its Qwen4 architecture, as open weights. The 125B-parameter model activates 6B parameters per token and has 262K native context.

RELEASE1mo ago
Perceptron releases Isaac 0.5 open weights for robot control

Perceptron released weights, inference code, and training details for Isaac 0.5, an embodied model for video perception, reasoning, and robot control. The 36B dynamic-MoE model was trained on 1 million hours of video.

RELEASE1mo ago
Z.ai releases GLM-5.3-Flash, identifies it as Ox Alpha

Z.ai identified the formerly anonymous Ox Alpha as GLM-5.3-Flash and released it under an MIT license. The native multimodal model has 320B total parameters, 18B active parameters, and a 1M-token context window.

RELEASE1mo ago
Perplexity launches Portable Computer on DGX Spark with a 27B model

Perplexity’s Portable Computer runs its orchestrator, subagents, and harness locally on NVIDIA DGX Spark with a post-trained 27B model. Frontier-model escalation requires user approval and flags PII before text is sent externally.

NEWS1mo ago
OpenRouter says Ox Alpha reaches 8 trillion daily tokens

OpenRouter says its free coding and sustained-agent model reached 8 trillion daily tokens within five days of launch. OpenCode separately reported processing 26 trillion Ox Alpha tokens over four days.

NEWS1mo ago
Qwen 3.8 27B reaches 3,200 TPM at 262K context on two RTX 3090s

Community tests report Qwen 3.8 27B handling coding, OCR, and long-context workloads locally. One vLLM setup reached 3,200 tokens per minute at 262K context on two RTX 3090s without NVLink.

NEWS1mo ago
Marin starts training open 535B-A23B model on 18.75T tokens

Marin has begun an open training run for a 535B-parameter mixture-of-experts model with 23B active parameters. The team plans to train on 18.75T tokens across 11 GB200 NVL72 systems over about three months.

WORKFLOW1mo ago
Users report Qwen 3.8 27B agents vary sharply by harness

Local users report that Qwen 3.8 27B agent behavior changes substantially with the harness, quantization, context settings, and hardware. One 3-bit MacBook Air run at 57K context took 63 hours.

NEWS1mo ago
LocalLLaMA post reports Qwen 3.8 27B reasoning loops caused most errors

A LocalLLaMA user reports reasoning loops caused most errors in a 2,483-task test of Qwen 3.8 27B. The report says 3–9% of inputs drove most failures because reasoning often did not terminate.

RELEASE1mo ago
Qwen releases Qwen3.8 27B multimodal model under Apache 2.0

Qwen released its open-weight Qwen3.8 27B vision-language model with 262K native context and adjustable reasoning. In a 484-sample test, enabled-thinking scores fell from above 92% through 64K to 74.3–81.8% at 128K.

RELEASE1mo ago
Qwen 3.8 Max 2.4T open weights ship as text-only, users say

LocalLLaMA users said Qwen 3.8 Max 2.4T open weights are text-only while the API keeps vision support. A linked Qwen3.8-27B ModelScope page reportedly returned 404 before release.

RELEASE2mo ago
Meta releases Muse Glimmer 30B as an Apache 2.0 open-weight local agent model

Meta released Muse Glimmer, a 30B Apache 2.0 dense model for local agent workflows. Reports cite 4-bit builds under 20GB, vision input, function calling, 131K context, and day-0 support in Hugging Face, vLLM, SGLang, Ollama, and MLX.

RELEASE2mo ago
Magnitude launches open-source offline coding agent for local models

Magnitude launched an open-source terminal coding agent that runs local models on-device without API keys. Its launch post says it profiles hardware and can use shell, file-editing, script, and skills tools.

RELEASE2mo ago
Qwen3.8-Max launches on OpenRouter with 1M-token context

Alibaba's Qwen3.8-Max is live on Venice and OpenRouter while open weights are still described as coming soon. Reports cite a 2.4T-parameter model with strong Vals and vision benchmark results.

RELEASE2mo ago
MiniMax H3 releases Hugging Face weights and fal video endpoints

MiniMax H3 now has Hugging Face weights, fal endpoints, AI Toolkit LoRA support, and reported single-RTX-5090 local runs. MiniMax also said deployment in the US, EU, UK, and South Korea is available through formal authorization.

RELEASE2mo ago
MiniMax releases H3 open weights with day-zero vLLM-Omni support

MiniMax released H3 weights on Hugging Face for text-to-video, image-to-video, reference-to-video, and editing workflows. vLLM-Omni, ComfyUI, SGLang Diffusion, and fal added support at launch.

RELEASE2mo ago
Alibaba says Qwen3.8-Max open weights ship next week

Alibaba said Qwen3.8-Max left preview as a 2.4T-parameter MoE with 95B active parameters and $2/$6 per million-token pricing. Arena placed it on the Frontend Code Arena cost-performance frontier.

RELEASE2mo ago
DeepSeek releases V4 Flash 0731 as MIT-licensed open weights

DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.

RELEASE2mo ago
Thinking Machines releases Inkling-Small 276B open-weight MoE

Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.

RELEASE2mo ago
LightOn releases mDenseOn and mLateOn retrieval models for 8 languages

LightOn released mDenseOn and mLateOn, which translate an open English retrieval recipe into 8 languages and ship models, data, and training code. The release includes 2.8B training pairs and 16.3M fine-tuning samples.

RELEASE2mo ago
MiniMax launches H3 for video apps and APIs

MiniMax announced H3 for app and API access, with OpenRouter, Vercel, Pika, fal, Runway, and Venice integrations. Artificial Analysis ranked it No. 1 for video editing, and MiniMax said open weights are planned.

NEWS2mo ago
Kimi K3 report details RL distillation and FlashKDA infrastructure

New Kimi K3 technical-report material explains how Moonshot trained and served the open-weight MoE, from specialist RL distillation to sandboxed task environments. Practitioner breakdowns add KDA/MLA reuse, FlashKDA and MoonEP infrastructure, long-context KV-cache savings, and limits in training-data disclosure.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.