Skip to content

Aleph Alpha Kolibri Ships as an Open-Weight English-German Model Under Apache 2.0

Aleph Alpha Kolibri is an open-weight English-German model with 78B parameters, about 3B active per token, released with full weights under Apache 2.0.

Kolibri logo, a stylized bird mark with the wordmark, in white on a green gradient background
Kolibri · Credit: Aleph Alpha

Aleph Alpha Kolibri arrives as an open-weight model that reasons in German and English, with 78 billion parameters of which about 3 billion are active for each token. The full weights sit on Hugging Face under the Apache 2.0 license.

The technical post describes a mixture-of-experts Transformer aimed at regulated work in public administration, industry and aerospace. Aleph Alpha lists a context window of up to 1 million tokens, yet the post's own model table gives 262,144 tokens as the longest length actually trained. On the company's benchmark table, Aleph Alpha Kolibri scores 96.9 on AIME 2025 math against 91.7 for Nvidia's Nemotron 3 Super, a model Aleph Alpha says carries up to four times the active parameters, and 84.3 on GPQA Diamond against 83.4 for Qwen3.6-35B-A3B.

Hallucination results are less flattering. Aleph Alpha Kolibri abstains instead of answering wrongly on 44 percent of AA-Omniscience items, up from 15 percent for its predecessor, but Qwen3.6-35B-A3B reaches 56.7 on the same measure. Every figure here is Aleph Alpha's own.

German is where the model is meant to differ. More than a fifth of the pre-training tokens are German, and a new tokenizer called UniBPE averages 4.90 bytes of German web text per token, against 4.35 for GPT-5's tokenizer in the company's comparison. English compression is a shade lower, at 4.58 against 4.67. Fewer tokens per document means lower serving cost for the long, compound-heavy prose that German administrations produce.

The size was a cost decision. In its tests a 123-billion-parameter variant handled three concurrent long-context queries on two H100 GPUs, while the 78-billion model handled 18 and decoded 28 percent faster, Aleph Alpha says. Pre-training ran 21 days on 768 B200 GPUs in Germany and Finland over 20 trillion tokens. Kolibri Origin, a 30-billion-parameter validation run with a much shorter context window, was never released publicly.

For buyers, the compliance paperwork may matter as much as the scores. Aleph Alpha says the technical report documents measures for the EU AI Act, copyright and data protection, including screening all training data against a blocklist of more than 4.5 million URLs. Open weights also let outside teams test the efficiency and grounding claims of Aleph Alpha Kolibri on their own hardware. The company, which agreed in September to combine with Cohere, says it keeps operating independently until that deal closes, and the transaction remains subject to regulatory approval.

Share this story

Kenji Sato

Kenji Sato edits techshooked's coverage of artificial intelligence and emerging technology, following the path from research to production systems. His standard is anti-hype: ask what a model actually does, what data trained it, how it fails in practice, and whether a benchmark measures what the marketing says it does.