Aleph Alpha Launches Kolibri-1 Reasoning Model
German startup Aleph Alpha has launched Kolibri-1, an open-weight reasoning model designed to give European enterprises high-performance local processing with a massive context window.

Heidelberg-based AI firm Aleph Alpha has launched Kolibri-1, an open-weight mixture-of-experts reasoning model optimized for German and English. Released under an Apache 2.0 license, the model features 78.1 billion total parameters but activates only 3.46 billion parameters per token. This sparse architecture relies on 384 routed experts per layer, selecting six active experts and one shared expert for each token. The model was trained from scratch on 20 trillion tokens using 768 NVIDIA B200 GPUs located in Germany and Finland.
Kolibri-1 stands out for its massive context capabilities, supporting a native window of 262,144 tokens that has been validated up to 1,048,576 tokens without requiring position scaling techniques. To manage this long context efficiently, the model uses a four-to-one ratio of sliding-window attention, which inspects the previous 512 tokens, to global attention layers. In benchmark evaluations, Kolibri-1 achieved a score of 96.9 on AIME 2025 and 84.3 on GPQA Diamond EN, outperforming comparable open mixture-of-experts models.
For practitioners, Kolibri-1 offers a powerful, self-hosted alternative for applications requiring strict data sovereignty and complex German-language reasoning. The model ships with FP8 weights, requiring approximately 78 GB of storage, which allows it to run on a single NVIDIA H200, B200, or B300 GPU. It can be served via a vLLM plugin, utilizing an FP8 key-value cache. Additionally, the model introduces the UniBPE tokenizer and the Merlin-Arthur reinforcement learning protocol, which helps control hallucinations during long-context retrieval and tool-use tasks.
This is our own summary of reporting by AlphaSignal


