Aleph Alpha's Kolibri Brings a Sovereign 78B Open-Weight Model to Europe
Aleph Alpha released Kolibri-1 on German Unity Day, an Apache 2.0 open-weight model with 78 billion total parameters and about 3.46 billion active per token, tuned for German and English. It leads its weight class on reasoning benchmarks and ships FP8 weights that self-host on two A100s or one H200, weeks after Aleph Alpha signed a merger agreement with Cohere.
On this page
A German-English MoE under the Apache license
Aleph Alpha, the Heidelberg company that has spent years positioning itself as Europe's sovereign alternative to US frontier labs, released Kolibri-1 on 3 October, choosing German Unity Day for the launch. The model is a Mixture-of-Experts transformer with 78 billion total parameters and only about 3.46 billion active per token, covering German and English with a German-tailored tokenizer, and the weights are on Hugging Face under the Apache 2.0 license.
The license choice is the headline for anyone who runs models locally. Aleph Alpha's earlier releases, such as Pharia-1-LLM, used the more restrictive Open Aleph License, which limited commercial use; Kolibri's weights and configuration ship under Apache 2.0, the same permissive terms as Mistral's and Qwen's open-weight lines, with only the training code and methods retained. That makes Kolibri the first genuinely self-hostable sovereign European model at this scale, and a direct answer to the question of what public administrations should run when the answer cannot be a foreign API.
The benchmarks, and the gaps
The technical report puts Kolibri at the front of its class: among MoE models with roughly 3 billion active parameters, it claims the lead in most categories, scoring 96.9 on AIME 2025 in English, 84.3 on GPQA Diamond, 85.9 on LiveCodeBench v6, and 66.4 on SWE-Bench Verified, with long-context results holding up at a million tokens on RULER. Native context is 262,000 tokens, extendable to about a million, and the model was trained on 20 trillion tokens across 768 B200 GPUs.
The card is also candid about the gaps: TerminalBench 2.1 at 27.7 and Tau3-Bench Banking at 38.1 are weak spots where larger models score higher. Deployment needs are modest by frontier standards, FP8 weights with a footprint around 78 GB served through vLLM on two A100 80GB cards, two H100s, or a single H200 or B200, which places self-hosting within reach of a serious department rather than only a hyperscaler.
Sovereignty, self-hosting, and the Cohere merger
The positioning matters as much as the numbers. Aleph Alpha pitches Kolibri for public administration and regulated industry, naming multi-step reasoning, coding, retrieval, structured extraction, and tool calling as intended workloads under human review, and the tech report frames it as a sovereign European model on the Pareto frontier of capability and controllability. The launch lands weeks after Aleph Alpha signed a definitive agreement with Canada's Cohere to form what the companies call the first transatlantic sovereign AI solution, with the combined business operating as Cohere from Berlin.
For the local AI community, Kolibri is worth downloading for three reasons. It is permissively licensed at a scale that fits one GPU, it is explicitly built for self-hosted rather than API-first use, and its German-first tokenizer makes it the strongest open option for a language most frontier labs treat as an afterthought. Whether the sovereign promise survives the Cohere integration is the open question; the weights, at least, are out, and Apache 2.0 means they stay out.