ESC
其他 8 分钟阅读

Show HN: Germany's new sovereign AI model Kolibri

Show HN: Germany's new sovereign AI model Kolibri

来源:Hacker News

In a normal (dense) model, every token goes through every parameter. In a mixture of experts (MoE), each layer has a crowd of small sub-networks called experts and a router that picks a few of them for each token. Kolibri has 50 layers, each with 384 experts plus 1 shared expert that every token goes through, and its router sends each token to 6 of the 384. That’s how 78.1 billion parameters turn into 3.46 billion of actual work per token.

In 2024 I gave a talk called Why Small Language Models are the future, and I argued for “smaller language models with fewer parameters and fewer places things can go wrong that require lesser compute.” My analogy was a doctor who has read every medical book in the world against a specialist in hematology: go to the first one with a blood condition and “they may not get it right cuz they know too much.” A mixture of experts puts a hospital full of specialists inside one model, and the router is the receptionist who sends each token to the right 6.

The analogy breaks in 2 places though. The experts aren’t neat topics like “German law”: when researchers look inside MoE models they mostly find experts for patterns of tokens, like punctuation or proper nouns, not subjects a person would pick. And the hospital has to keep all 384 specialists on staff even if you only see 6, so Kolibri computes like a 3.5 billion parameter model but needs the memory of a 78 billion parameter one. The model card says it plainly: “the full model must be held in memory even though only part of it is active at any time.”

A model doesn’t read letters or words, it reads tokens: chunks of text from a fixed vocabulary, picked when the tokenizer is trained. German glues words together into long compound words, and a tokenizer that learned mostly from English chops them into pieces. Here’s the German name of the Federal Constitutional Court, split by the tokenizer GPT-4o and GPT-5 use (o200k_base, through OpenAI’s tiktoken), and by Kolibri’s:

o200k_base (GPT-5): Bund | es | ver | fass | ungs | gericht 6 tokens Kolibri: Bundes | verfassungsgericht 2 tokens Kolibri’s tokenizer has 128,000 tokens, trained with a new algorithm Aleph Alpha calls UniBPE: it keeps the bottom-up merging of byte-pair encoding (BPE) and picks each merge with a different scoring rule (the Unigram objective), which respects how German builds words. The report says it needs 11.2% fewer tokens for German text than GPT-5’s tokenizer, the best of the 9 others they measured.

I wanted to see that for myself, so I ran 6 tokenizers over all of the Basic Law for the Federal Republic of Germany, the German constitution (185 KB of very German legal text), and over its official English translation:

On legal German, Kolibri needed 15% fewer tokens than GPT-5’s tokenizer, even more than Aleph Alpha’s own 11.2%, and in English it tied with it. Wild. Fewer tokens means fewer steps to read or write the same German text, and more German fits in the same context window. (I counted each one with its own tokenizer.json through Hugging Face’s tokenizers library, except o200k_base, which I counted with tiktoken.)

40 of Kolibri’s 50 layers use sliding-window attention: each token only looks at the 512 tokens before it. Every 5th layer looks at everything before it. It’s like reading a long contract while mostly paying attention to the sentence you’re on, and every few pages stopping to think about all of it, and it’s what keeps a 1 million token context affordable.

There’s a clever detail in there too. Only the sliding-window layers know where a token sits (through rotary position embeddings), and the full-attention layers don’t, so the context stretches past the 262,144 tokens it was trained on without any extra position tricks. Aleph Alpha validated it up to 1,048,576.

In my talk Unlocking Value with AI Today I called finite context one of “the big three” problems of generative AI, next to hallucination and the knowledge cutoff, and Kolibri goes after all 3. At 1 million tokens, on the RULER long-context test, Kolibri’s base model scores 63.2, against 57.5 for Qwen3.5 35B-A3B’s base model.

Reasoning models think before they answer, and even on German prompts they mostly think in English. Aleph Alpha posted about this on 24 September and wrote it up as Through the Valley of Tears: they generated about 800,000 German reasoning examples, and found that a little German reasoning data is worse than none. Their model’s German math score dropped from 70.2 to 48.3, because its German thoughts kept going around in circles and never finished, and it only climbed back (to 67.3) with a lot more German data.

Kolibri got the lot more. It reasons in German on German prompts, and its German math scores are the best of the models with about 3 billion active parameters: 87.5 on the American Invitational Mathematics Examination (AIME) 2025 in German, against 84.4 for the next best, NVIDIA’s Nemotron 3 Nano.

Hallucination was number 1 of my big three, and the fix in that talk was retrieval augmented generation (RAG): you look up good, authoritative information and put it in the prompt. RAG only works if the model admits when the documents don’t have the answer, though, and that’s what Aleph Alpha trained for, with their own method, the Merlin-Arthur protocol.

It works like a game with 3 players. Arthur is the model, and he gets a question with parts of the supporting document hidden. Merlin hides parts so that the correct answer gets easier to find, and Arthur is trained to answer those. Morgana hides the evidence the answer depends on, and Arthur is trained to say he doesn’t know. Arthur never knows which of the 2 he’s facing, so the only way to win is to actually check whether the evidence in front of him supports an answer.

It shows. On Artificial Analysis’s Omniscience test, when Kolibri didn’t know an answer, it said so (or gave a partial answer) 44% of the time instead of making one up. Qwen3.5 35B-A3B did that 11.1% of the time and GPT-OSS 120B 23.7%. Of all the mixture-of-experts models Aleph Alpha compared, only Qwen3.6 35B-A3B did better, at 56.7%.

Every request can set reasoning_effort to none, low, medium or high, so a quick lookup answers right away and a hard question gets a long think, from the same model on the same server.

These are the rows where Kolibri leads the open models of its size in Aleph Alpha’s evaluation, which runs every model through the same setup with the sampling settings its makers recommend:

The math is the standout: on AIME 2025 and 2026 in English it beats every MoE model in the comparison, including the ones with 12 billion active parameters, and only the dense Qwen3.8 27B scores higher. The company-documents row is 5 tests Aleph Alpha built from customer-like work in semiconductors, the German public sector, aerospace, an automotive supplier and industrial drives, run over documents and questions the model never saw in training.

Aleph Alpha publishes its weakest rows right next to its best ones in the model card, so here they are:

You need about 78 GB of GPU memory: 2 NVIDIA A100s or H100s with 80 GB each at the least, or a single H200, B200 or B300. I haven’t run the model itself yet, since it needs a data-center GPU and nobody hosts it so far, so this is straight from the model card. First the plugin, which installs the vLLM version it supports:

`pip install ‘aleph-alpha-inference>=1’

vllm serve Aleph-Alpha/Kolibri-1 –kv-cache-dtype fp8
–reasoning-parser kolibri1
–tool-call-parser kolibri1
–enable-auto-tool-choice` That gives you an OpenAI-compatible server, so any OpenAI client talks to it, and the reasoning effort goes through the chat template:

`# from the Kolibri model card (trimmed) from openai import OpenAI

client = OpenAI(base_url=“http://localhost:8000/v1”, api_key=“EMPTY”) response = client.chat.completions.create( model=“Aleph-Alpha/Kolibri-1”, messages=[{“role”: “user”, “content”: “Erkläre kurz, was ein Mixture-of-Experts-Modell ist.”}], extra_body={“chat_template_kwargs”: {“reasoning_effort”: “high”, “enable_thinking”: True}}, ) print(response.choices[0].message.content)The model card recommendstemperature=1.0, top_p=0.97andtop_k=128`, and contexts of at most 262,144 tokens for anything latency-sensitive. For the full million, serve it with 2 more flags:

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \ --max-model-len 1048576 \ --hf-overrides '{"max_position_embeddings": 1048576}'

When to use Kolibri

Kolibri is the pick when German text and your own hardware both matter: a public authority, a bank, a manufacturer or an aerospace supplier that has to keep its documents in house, wants answers in German that reason in German, and would rather hear “I don’t know” than a confident wrong answer. RAG over long German documents (laws, contracts, manuals) plays to every strength above: the tokenizer, the 1 million token context and the abstention.

That’s the kind of project I’m working on right now, with Prof. Dr. Heinrich Audebert, who heads neurology at Campus Benjamin Franklin, one of the Charité’s hospitals in Berlin. Today, a patient with a neurological complaint goes to their general practitioner (GP), and the GP has to see them, which is very demanding for a busy practice. In what we’re building, the patient sits down at a computer in the GP’s practice, an AI avatar takes them through a battery of tests, and it grades how urgent their symptoms are: a referral to a specialist right away, or they can wait a bit. The conversations are in German, they’re about people’s health, and the model has to run where we control it, so we’re thinking of using Kolibri for it.

It’s the wrong pick for a coding agent, where Qwen3.6 35B-A3B leads, for questions the model has to answer from memory, for any language besides German and English, and for anyone who can’t spare 78 GB of GPU memory. For that last case, the small specialized model I argued for in 2024 is still the answer: for my podcast search I fine-tuned Mistral 7B on my Apple silicon laptop instead of paying for GPT-4o.