Aleph Alpha's Kolibri-1 Brings Open German-English Reasoning to a Single GPU
Aleph Alpha ships a 78B Mixture-of-Experts model with 3.46B active parameters, a 1M-token context, and Apache 2.0 weights focused on German and English.
- Aleph Alpha released Kolibri-1, a 78B MoE model with 3.46B active parameters under Apache 2.0.
- Context window validated to 1,048,576 tokens; native trained length is 262,144 tokens with hybrid sliding-window attention.
- Scores 96.9 AIME 2025, 84.3 GPQA Diamond EN, 66.4 SWE-Bench Verified, 85.9 LiveCodeBench v6.
- FP8 weights total ~78 GB; runs on a single H200/B200/B300 or 2x H100.
- Trained from scratch on 20T tokens across 768 B200 GPUs, with custom UniBPE tokenizer for German morphology.
- Uses novel Merlin-Arthur RL protocol to reduce hallucinations in RAG workflows; strong German-language industry benchmarks.
Kolibri-1 packs German-English reasoning into a sparse 78B model
German AI company Aleph Alpha has released Kolibri-1, an open-weight reasoning model focused on English and German. Its mixture-of-experts architecture contains 78.1 billion parameters but activates about 3.46 billion for each token, reducing inference computation while retaining a larger model’s capacity.
Kolibri-1 gives European organizations an Apache 2.0 model they can run on their own infrastructure, with minimum requirements of 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300. Aleph Alpha controlled the data pipeline, architecture, tokenizer, training infrastructure, post-training and evaluation. Training ran in Germany and Finland, with development aimed at the EU General-Purpose AI Code of Practice, the EU AI Act and GDPR.
78 billion stored, 3.46 billion active
| Component | Specification |
|---|---|
| Architecture | 50-layer mixture-of-experts transformer |
| Parameters | 78.1B total, approximately 3.46B active per token |
| Experts | 384 candidates per layer, with six routed experts and one shared expert active |
| Attention | 512-token sliding-window attention, with full attention every fifth layer |
| Native context | 262,144 tokens |
| Maximum context | 1,048,576 tokens through validated extrapolation |
| Weights | FP8, approximately 78 GB |
| License | Apache 2.0 |
Mixture-of-experts routing sends each token through a small subset of the available feed-forward networks. This limits the matrix multiplication required for each token. The full 78.1 billion parameters still determine weight storage and memory requirements, while context length, batch size and concurrency add substantial key-value cache overhead.
A million-token ceiling, with caveats
Kolibri-1 was trained natively at up to 262,144 tokens. Aleph Alpha validated extrapolation to 1,048,576 tokens, but recommends staying at or below the native length for efficient serving on complex tasks. Teams considering the larger window should measure retrieval accuracy, latency and memory use on representative documents.
The approximately 78 GB FP8 checkpoint fits on one H200, B200 or B300 according to Aleph Alpha. Two 80 GB A100s or two H100s can also host it. Those figures describe weight capacity; production sizing must also account for the runtime, key-value cache, batching and tool-calling workload.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.