Aleph Alpha releases Kolibri, an open-weight 78B MoE model

On the Day of German Reunification, Aleph Alpha released Kolibri, an English-German Mixture-of-Experts Transformer with 78B total parameters and 3B active per token. It supports context lengths of up to 1M tokens. The full weights are on Hugging Face under the Apache 2.0 license.
The company positions Kolibri as a specialized model for sovereign, mission-critical work in regulated areas including public administration, industrials and aerospace. It was specialized for German, reasoning, math, agentic behavior and other capabilities customers need in production. Aleph Alpha says it built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control, and that it owns the whole pipeline from data curation through pre- and post-training. It says it built Kolibri with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in mind. The small size, it says, lets customers run the model on-premise without sending internal data to third-party inference services.
On quality, the company claims Kolibri sits on the Pareto frontier of quality versus serving cost for both English and German: none of the compared models gives more quality at the same cost, or the same quality at lower cost. Across math, coding, grounding and long-context tasks, it says Kolibri matches models with up to four times its active parameter count, such as Nemotron 3 Super. Because public benchmarks miss sector needs, Aleph Alpha built internal evaluation suites for verticals such as the German public sector, aviation, manufacturing and automotive, with paired synthetic training environments, and says it never trained on customer data. For grounding, it used abstention data and its Merlin-Arthur protocol, so the model is trained to say "I don't know" when the answer is not in the context. The model was also taught to reason at four different effort levels.
On language, Aleph Alpha built a bilingual German/English tokenizer and used organic German data throughout, so 21.3% of pre-training tokens are German (about 4.3T tokens, against roughly 62% English and 14% code). Translated data was used sparingly, 6% overall.
Much of the post describes what the company calls its Model Factory, a training pipeline implemented as code, with each proposed change triggering a small end-to-end run and runs executed as GitHub Actions workflows. The pipeline was first validated on Kolibri Origin, a 30B total, 3B active model with a 65k-token context window. Work on the pipeline began in January; Kolibri Origin finished pre-training on 11 June and Kolibri on 11 September. In those three months the team went from 30B to 78B parameters, from 65k to up to 1M context, and from 7.5T to 20T training tokens. To get 20T tokens the pipeline processed over 200T tokens of raw data. The team also changed the attention design, tripled the number of experts, increased sparsity, replaced the routing algorithm, improved post-training data and more than doubled the number of environment tasks.
Architecture details: the post says 78B is 2.5 times the size of Kolibri Origin. Experiments with a model growing from 32B to 123B kept improving performance, but cost decided it: on two H100s a 123B model handles only 3 long-context 256k-token queries, while the 78B handles 18 concurrent requests and decodes 28% faster. Kolibri uses 384 smaller experts rather than fewer wide ones. Of 50 layers, only 10 process full context; the other 40 use a 512-token window.
Training ran on 768 B200 GPUs in three stages: 20T tokens of pre-training at 16k sequence length over 21 days, 3.44T tokens of mid-training at 64k, and 200B tokens of long-context adaptation at 256k, nearly 24T tokens in total, roughly three times what Kolibri Origin consumed. Over the 21 days of pre-training there were 38 unplanned interruptions, roughly one per 10,000 GPU-hours, from hardware faults or connection timeouts. The pipeline handled them automatically, restarting on different nodes from a checkpoint at most 250 steps back.
The authors are candid about setbacks: they stopped Kolibri Origin pre-training after a few trillion tokens and restarted from scratch because of a data-shuffling bug that escaped their tests, and some ablations had to be rerun. They call the gain over Kolibri Origin notable but the easy direction: more parameters, more data, a well-understood architecture family, still at small scale.
Key facts
- Kolibri is an English-German Mixture-of-Experts Transformer with 78B total and 3B active parameters, context up to 1M tokens, full weights on Hugging Face under Apache 2.0.
- Aleph Alpha claims Kolibri sits on the quality-versus-serving-cost Pareto frontier for English and German and matches models with up to four times its active parameters, such as Nemotron 3 Super.
- It was trained on 768 B200 GPUs: 20T pre-training tokens over 21 days, 3.44T mid-training tokens and 200B long-context tokens, with 21.3% of pre-training tokens in German.
- The 78B size was chosen over 123B because on two H100s the larger model handles only 3 long-context 256k-token queries versus 18 concurrent requests, and decodes 28% slower in the comparison.
- The company says Kolibri was built and trained in Germany and Finland with no foreign control, aimed at regulated sectors such as public administration, industrials and aerospace.
Why it matters
Kolibri is an open-weight European model sold on sovereignty: German-built, trained on German and Finnish infrastructure, with no foreign control, and licensed Apache 2.0. Only 3B of its 78B parameters are active per token, which the company ties to lower serving cost and on-premise deployment. The post also shows the speed of the team's process: Kolibri Origin and Kolibri finished pre-training three months apart (11 June and 11 September), with the larger model trained on 20T tokens against 7.5T.
Who it affects
The company aims Kolibri at customers in regulated areas, naming public administration, industrials and aerospace, and sectors such as the German public sector, aviation, manufacturing and the automotive industry. Anyone who needs a German-capable model they can run on their own hardware, without sending internal data to third-party inference services, is the target audience. With weights on Hugging Face under Apache 2.0, others can download and use it too.
How to use it
The weights can be downloaded in full from Hugging Face and used under Apache 2.0. The model supports controllable reasoning effort, with four effort levels, to trade cost and latency against answer quality. It is trained to abstain with "I don't know" when the answer is not in the supplied context, which suits document-grounded use. The post links a tech report for full details; it gives no pricing.
How solid is it
This is the company's own announcement, and every performance and sovereignty claim in it is self-reported. The visible text gives no benchmark scores, only the Pareto-frontier claim and the claim of matching models with up to four times the active parameters; the numbers are in the linked tech report. It names Nemotron 3 Super as one comparison and does not say which other models were compared. Concrete training details (GPU count, token counts per stage, 38 interruptions, the 384 experts, the layer layout) are specific and internally consistent. The authors also admit a restart caused by a data-shuffling bug.
Risks and caveats
No independent verification of the benchmark or sovereignty claims is given. The EU AI Act, GPAI Code of Practice and GDPR are described as design intent ("in mind"), not a certification. The authors themselves say the gain over Kolibri Origin is the easy direction, with more parameters and data on a well-understood architecture at small scale. The Pareto claim covers the models the company chose to compare. No hardware requirements for running Kolibri are given, and the post gives no calendar date for the release beyond the Day of German Reunification.
“Kolibri is a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace.”
— Aleph Alpha, Kolibri announcement