News

Aleph Alpha releases Kolibri for German and English workflows

In brief

Kolibri adds a customer-controlled bilingual model option. Its 3.46B active parameters come with about 78 GB of FP8 weights, plus memory for request state.

4 min read

Sources
A linocut hummingbird visits coral flowers in a large green terraced garden.
Original AI-generated conceptual illustration: a hummingbird visiting part of a larger garden symbolizes selected computation. It is not a Kolibri architecture diagram or measured result.
On this page3 sections

Aleph Alpha has released Kolibri, a German- and English-focused language model that companies can run on their own infrastructure. The weights became available on October 3; the company's October 5 product announcement positions it for document processing, retrieval and agent workflows. For a team supporting services across both languages, that makes it a candidate for searching internal material and drafting answers without sending the entire workflow to a hosted model.

The serving details are more revealing than the small active-parameter number. Kolibri activates 3.46 billion parameters for each token, but its model card lists about 78 GB of FP8 weights. The distinction matters before anyone reserves a GPU or promises interactive response times.

Small active compute, a larger memory footprint

Kolibri is a mixture-of-experts model. Instead of using every expert for every token, it routes the token through a selected subset. That reduces the computation performed at that step. The experts that were not selected still need to be available for later tokens, so a low active count does not mean the rest of the model disappears from memory.

The published minimum configurations include one H200, B200 or B300, or two A100 80 GB or H100 SXM5 accelerators. Those configurations describe a documented deployment footprint. They do not tell an SRE how many simultaneous support investigations will meet the service's response-time target. Request length, generated output and concurrency remain part of that capacity question.

Two separate Kolibri serving considerations: approximately 78 GB of FP8 weights stay resident, while 3.46 billion parameters are active for each token. Request state needs additional capacity.
The active-parameter count describes computation per token. The weight footprint describes stored model data; request state adds another capacity consideration. Conceptual explanation, not a hardware benchmark.

The context limit has a similar distinction. The card identifies 262,144 tokens as the native context and validates extended support up to 1,048,576. Aleph Alpha recommends staying at or below the native length for complex tasks and deployments sensitive to latency or throughput. The million-token ceiling offers an option; it is not the provider's default advice for every production request.

What the serving comparison establishes

Aleph Alpha's technical report compares quality and serving cost using synthetic prompts and high-concurrency decoding on a node with eight B200 accelerators. The post-trained throughput comparison uses a 16K context. It is evidence about that experiment, with each model assigned its fastest tested parallel layout, rather than a measurement of a one-GPU assistant responding to an engineer.

A team buying capacity for a bilingual support assistant needs evidence at its own request lengths and concurrency, including the first requests after a cold restart. Our AI storage and recovery analysis makes the same distinction between a tested configuration and the workload expected to use it. The eight-GPU graph can inform that comparison; it cannot replace it.

Consider a hypothetical support assistant that searches German runbooks and English incident notes. A useful comparison would give Kolibri and the current model the same questions and source material, then check whether the answers cite the relevant passage, preserve its conditions and leave the engineer less correction work. Run that task at the expected number of simultaneous sessions and record completed answers alongside response time. These are proposed evaluation steps, not results from a Kolibri deployment.

If the workflow mainly needs other languages, or the available accelerator cannot hold the model and its request state, Kolibri's headline efficiency does not resolve the mismatch. A smaller model or a hosted service may be the more feasible comparison. For a bilingual workload that fits, the release creates a concrete option to test.

Integrating the release

Serving uses Aleph Alpha's vLLM plug-in package and its supported vLLM version, with Kolibri-specific reasoning and tool-call parsers. The release blog describes four reasoning settings, from none through high. Changing that setting changes the work requested from the model, so keep it recorded when comparing behavior.

The model card applies Apache 2.0 to the published weights and configuration files and excludes other artifacts from that grant. It also describes human-reviewed outputs and validated tool results as the intended integration. Self-hosting supplies control over where the model runs; the surrounding retrieval, review and execution workflow still determines what an operational assistant can safely do.

Kolibri's useful question is therefore specific: can a customer-controlled German/English model deliver the needed answer, at the actual request size and load? The release supplies the model and serving documentation. The decision comes from the workload that will use them.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question