Aleph Alpha releases Kolibri with a 1M-token context window
Aleph Alpha released Kolibri, an Apache 2.0 open-weight MoE with a 1M-token context window. The model has 78B total parameters and 3.46B active parameters, plus reasoning, tool calling and German-focused training.

TL;DR
- Aleph Alpha shipped Kolibri under Apache 2.0: 78B total parameters, 3.46B active per token and a 1M-token context window, according to the release overview.
- Reasoning and tool calling are built in, as the launch coverage reports.
- German accounts for more than a fifth of pre-training tokens, with a dedicated bilingual tokenizer described in the announcement coverage.
- Cohere hinted at joint development in its reaction to the release, without naming a product.
The 1M-token window comes with refreshingly specific fine print: Aleph Alpha recommends 256K or less for complex tasks and serving efficiency. Its training write-up records 38 unplanned pre-training interruptions, all recovered without manual intervention.
Open weights
Aleph Alpha released Kolibri on German Unity Day, October 3, 2026. Full weights are available in the Hugging Face repository, and the official serving package also lists a BF16 checkpoint.
Cohere called the model “really good” and added: “you're going to love what we build together.”
MoE architecture
Kolibri's sparse execution lowers per-token compute, but the full model stays in memory. Aleph Alpha's model card puts the FP8 weight footprint at roughly 78 GB.
Kolibri combines hybrid attention and fine-grained experts:
- Experts: 50 MoE layers, each with 384 routed experts, six selected per token, plus one shared expert.
- Attention: 40 sliding-window layers covering 512 preceding tokens, plus 10 full-context layers.
- Heads: 48 query heads and four KV heads.
- Precision: FP8 weights in 128×128 blocks; embeddings, output head, norms and router remain BF16.
- Listed minimum hardware: Two A100 80 GB GPUs, two H100 SXM5 GPUs, or a single H200, B200 or B300.
Million-token context
Kolibri's native context length is 262,144 tokens. Aleph Alpha validated quality and serving efficiency up to 1,048,576 tokens; positional encoding appears only in the sliding-window layers, allowing extension without position scaling.
Training progressed through three sequence lengths:
- Pre-training: 16,384-token sequences over 20T tokens.
- Mid-training: 65,536-token sequences over 3.44T tokens.
- Long-context extension: 262,144-token sequences over approximately 201B tokens.
Beyond the native window, the published vLLM example adds two configuration overrides:
Reasoning and tool calls
Version 1.0.0 of Aleph Alpha's serving plugin targets vLLM 0.29. It registers the model architecture and dedicated kolibri1 reasoning and tool-call parsers.
The documented serving command is:
Request controls and formats are documented in the model card:
- Reasoning effort:
none,low,mediumorhigh, passed throughchat_template_kwargs. - Thinking switch: On by default;
enable_thinking=falsedisables it. The reasoning parser respects the switch when separating reasoning from answer content. - Tool schemas: Functions go in the standard
toolsfield of an OpenAI-compatible chat request. Hermes-style calls become structured tool output and can be combined with reasoning. - Vendor-recommended sampling:
temperature=1.0,top_p=0.97,top_k=128.
Benchmarks
Aleph Alpha evaluated Kolibri at high reasoning effort, using eval-framework for most benchmarks and Harbor for SWE-Bench and TerminalBench. Its 30.6B comparison model, Kolibri Origin, was never publicly released.
The first-party results report these gains over Origin:
| Benchmark | Kolibri Origin | Kolibri | Gain |
| --- | ---: | ---: | ---: |
| AIME 2026, English | 81.5 | 96.0 | +14.5 points |
| GPQA Diamond, English | 68.1 | 84.3 | +16.2 points |
| LiveCodeBench v6 | 59.2 | 85.9 | +26.7 points |
| τ3-Bench Banking | 5.7 | 38.1 | +32.4 points |
Kolibri trails Qwen3.6-35B-A3B on three engineering-relevant rows in the same results:
- BFCL v4 overall: 61.4 versus 67.2, 5.8 points lower.
- LongBench Pro: 64.5 versus 70.8, 6.3 points lower.
- SWE-Bench Verified: 66.4 versus 73.8, 7.4 points lower.
Each model uses its documented context window and sampling parameters. On long-context benchmarks, an overlength prompt receives a score of zero.
German training
Aleph Alpha specialized Kolibri's 128,000-entry tokenizer for German word structure while retaining English support.
The German pre-training share is 21.3% in the launch post and 23.9% in the model card. The launch post says translation accounts for 6% overall, with an emphasis on organic German data because translations retain the cultural fingerprint of their source language.
European training and copyright
Teams in Germany developed Kolibri, and training ran on infrastructure in Germany and Finland, according to the launch write-up. Aleph Alpha says it designed the development pipeline around the EU AI Act, the General-Purpose AI Code of Practice and GDPR, with copyright law a particular focus.
The company also publishes a copyright-law due-diligence example in its 189-page technical report.
Grounded answers
Kolibri learned to abstain when supplied context does not support an answer through abstention data and Aleph Alpha's Merlin-Arthur protocol. On RGB Negative, an abstention evaluation, it scored 85.6 versus Origin's 73.9, +11.7 points, in the published results.
Aleph Alpha defines its intended systems as human-reviewed assistants and agentic workflows, with a person reviewing outputs before they are acted on. The model card places decision-support deployments on the advisory side, surfacing evidence and drafting options for human judgment.