Aleph Alpha Releases Kolibri-1, an Open German Reasoning Model With 1M-Token Context
Aleph Alpha drops a 78B mixture-of-experts reasoning model with 3.46B active parameters, 1M-token context, and Apache 2.0 weights, tuned for German-English sovereign deployments.
- Kolibri-1 is a 78B MoE with 3.46B active parameters, Apache 2.0, German and English focus.
- Context window validated up to 1,048,576 tokens, native 262,144, no position scaling tricks required.
- Scores 96.9 on AIME 2025 and 84.3 on GPQA Diamond EN, best among compared MoE models.
- Trained from scratch on 20T tokens using 768 NVIDIA B200 GPUs in Germany and Finland.
- Ships FP8 weights (~78 GB), runs on a single H200, B200 or B300; served via vLLM plugin.
- Introduces UniBPE tokenizer and Merlin-Arthur RL protocol for hallucination control (paper).
Aleph Alpha releases Kolibri-1 with sparse compute and a 1M-token option
Heidelberg-based Aleph Alpha has released Kolibri-1 model card, an open-weight mixture-of-experts reasoning model for German and English. It contains 78.1 billion parameters while activating about 3.46 billion for each token, supports configurable reasoning and tool calls, and carries an Apache 2.0 license for self-hosted deployment.
Kolibri-1 targets teams that need German-language reasoning, long-context retrieval and control over where prompts and model weights reside. Its sparse architecture reduces per-token computation, although the full parameter set still requires roughly 78 GB of weight storage.
Sparse compute, dense memory
| Specification | Kolibri-1 |
|---|---|
| Total parameters | 78,103,074,560 |
| Active per token | 3,457,573,120 |
| Languages | German and English |
| Transformer layers | 50 |
| Experts | 384 routed experts per layer, with six selected per token and one shared expert |
| Native context | 262,144 tokens |
| Validated extension | Up to 1,048,576 tokens |
| Weight format | FP8, with selected components in bfloat16 |
| License | Apache 2.0 |
The 22.6-to-1 ratio between total and active parameters explains the model’s unusual hardware profile. Each token uses only a fraction of the expert weights, which limits computation, while serving still requires access to the complete network.
Every transformer block uses a mixture-of-experts feed-forward layer. Attention follows a four-to-one pattern of sliding-window and full-attention layers: most layers inspect the previous 512 tokens, while periodic global layers can use the full context. This design limits attention work across most of the stack while preserving routes for long-range information.
Long context without position scaling
Aleph Alpha trained the model at 16,384 tokens, continued training at 65,536, and extended the native context to 262,144. The company reports successful quality and serving tests at 1,048,576 tokens.
Positional encoding appears only in the sliding-window attention layers, according to the model documentation. That architecture allows Aleph Alpha to extend the window without YaRN or RoPE scaling. Actual request limits still depend on the vLLM configuration, batch size and memory available for the key-value cache.
Weights use FP8 in 128-by-128 blocks, and activations are dynamically quantized. Embeddings, normalization layers, the language-model head and the expert router remain in bfloat16. Aleph Alpha evaluated the model with an FP8 key-value cache.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.