
Kolibri is a large language model built end to end in Europe. Aleph Alpha, from Heidelberg, curated the data, trained the model on GPUs in Germany and Finland, and released the weights openly on 3rd October 2026 under the Apache 2.0 licence. It focuses on 2 languages, German and English, and on cheap serving: of its 78 billion parameters, only about 3.5 billion work on each token.
It's aimed at regulated work: public administration, industry, aerospace. That's where "where was this built, and who controls it" matters as much as the benchmark score.

| Maker | Aleph Alpha GmbH, Heidelberg, with offices in Berlin, Bayreuth and Munich |
| Released | 3rd October 2026, open weights on Hugging Face |
| Architecture | Mixture-of-experts, 50 layers, 384 experts per layer (1 shared, 6 routed per token) |
| Parameters | 78B in total, 3.46B active per token |
| Languages | German and English, with a tokenizer built for German word structure |
| Context | 262,144 tokens native, validated up to 1,048,576 |
| Modes | Reasoning with effort levels (none, low, medium, high), tool calling |
| Knowledge cutoff | 18th June 2026 |
| Licence | Apache 2.0 |
| Compliance | Aleph Alpha signed the EU General-Purpose AI Code of Practice |
What makes it different
- Built in Europe, under European law. Trained on 768 NVIDIA B200 GPUs in Germany and Finland, by a German team that owns the whole pipeline from data to optimisation. The company says there's "no foreign control".
- German as a first language, not a translation. 23.9% of the pre-training data is German (per the model card), mostly original German text rather than machine translation.
- Cheap per token. Only 3.5B parameters are active per token, so it runs fast for its quality. The catch: all 78B must sit in memory.
- Long documents. Most attention layers look at nearby text, a few look across the whole context. That keeps a 1M-token context affordable.
- Transparent. The model card documents the training data mix, the compute and the energy: about 950 MWh for pre-training and long-context training.
For business people
The problem: many European organisations can't, or won't, send sensitive documents to a US model. The open alternatives mostly come from China or the US, and their training data is a black box.
What Kolibri offers: a capable model with a European chain of custody. You can run it on your own servers or a European cloud, inspect the weights, and point regulators to a documented training process.
Where it fits:
- Document work: drafting, summarising and extracting from long German or English files.
- Question answering over your own material, with retrieval.
- Agents: calling tools and APIs, with a person reviewing the output. Aleph Alpha positions it as an advisor, not as the component that decides.
How good is it? On Aleph Alpha's own overall scores, Kolibri reaches 75.5 in English and 70.8 in German, the best of the small-active-parameter models it compares, ahead of Qwen3.5 35B-A3B (74.7 and 69.8). Larger dense models still score higher: Qwen3.8 27B reaches 80.2 and 79.9. Aleph Alpha also reports 96.9 on the AIME 2025 maths test and 84.3 on GPQA Diamond.
The limits:
- 2 languages only. Deliberately: depth over breadth. No French or Spanish focus.
- Big hardware. About 78 GB of memory for the weights. The minimum is a single H200 or B200, or 2 A100 or H100 GPUs.
- The scores are the maker's own, on its own evaluation set.
For technical people
Kolibri needs Aleph Alpha's vLLM plugin:
pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
The server speaks the OpenAI API, so any OpenAI client works:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Aleph-Alpha/Kolibri-1",
messages=[{"role": "user", "content": "Fasse diesen Vertrag in 5 Punkten zusammen."}],
extra_body={"chat_template_kwargs": {"reasoning_effort": "medium"}},
)
print(response.choices[0].message.content)
- Precision: FP8 weights in 128×128 blocks, with an FP8 key-value cache. Embeddings, output head, norms and router stay in BF16. A BF16 version is published separately.
- Attention: 4 sliding-window layers for every global layer. Positions are encoded only in the sliding-window layers, which is why the context can stretch past its training length.
- Training: 20 trillion tokens of pre-training (62.5% English, 23.9% German, 13.6% code), plus mid-training, long-context training, supervised fine-tuning and reinforcement learning. Optimiser: Muon.
- Sampling: Aleph Alpha recommends temperature 1.0, top_p 0.97 and top_k 128.
- Long context: add
--max-model-len 1048576and the matching--hf-overridesto go past 262,144 tokens, at a cost in speed. - Tool calling: Hermes-style function calling through the standard
toolsfield, combinable with reasoning.
Value
This is the model I'd point to in a sovereignty conversation with a German customer: trained in Europe, documented, open weights, Apache 2.0, and strong in German. It's the model-side answer to a European inference provider like SovInfra (see SovInfra ).
For my own setup, it doesn't fit yet. My NicAI model runs on Qwen3.8-27B in Ollama, and Kolibri ships for vLLM on data-centre GPUs, not as a local Mac model. If someone publishes a quantised version that runs on Apple Silicon, it would be worth a test on German writing, where most open models are weakest.
Further Reading


