agentos.guide › blog
Kolibri AI is a free, open-weight model from Aleph Alpha in Germany, with 78.1 billion parameters, about 3.46 billion active per token, and native German and English.
People also search for it as Colibri, but the official name is Kolibri, which means hummingbird in German.
The question I get most is whether it can be a free, private brain for your agents.
It can, if you have the memory for it, so here is the quick-start, the honest hardware bill and exactly where it fits in an Agent OS.
Kolibri AI is the open-weight model Aleph Alpha released on 3 October 2026, and its official name is Kolibri 1.
Kolibri is the German word for a hummingbird, which fits a big model that only flaps a small part of itself for each word.
You will also see it written as Colibri, and the auto-captions on my own video even spelled it Calibri.
Here is the spec sheet straight from the official Hugging Face model card.
| Spec | Kolibri 1 (official model card) |
|---|---|
| Maker | Aleph Alpha, a German AI company based in Heidelberg. |
| Hugging Face repo | Aleph-Alpha/Kolibri-1 holds the FP8 weights, and Aleph-Alpha/Kolibri-1-BF16 holds full precision. |
| Released | It was released on 3 October 2026. |
| Licence | It uses Apache 2.0, and the repo is not gated. |
| Size | It has 78.1 billion total parameters, with about 3.46 billion active per token. |
| Languages | It speaks German and English natively. |
| Context | It reads 262,144 tokens natively, and Aleph Alpha validated it up to 1,048,576. |
| Knowledge cutoff | Its built-in knowledge stops at 18 June 2026. |
| Agent features | It has a reasoning mode with effort levels and Hermes-style tool calling. |
| Input and output | It is text only. |
In my video I rounded the active parameters to "3 billion", and the exact official figure is 3.46 billion.
I also said it handles up to 1 million tokens, which is right, but Aleph Alpha recommends staying at or under 262,144 tokens for speed and for complex tasks.
🔥 Want the exact local models in the Agent OS setup I run inside my Agent OS?
Inside the AI Profit Boardroom you get the installable Agent OS, step-by-step video tutorials, four coaching calls a week and 3,400+ members building real automations.
An Agent OS is just a set of agents, tools and memory that run your work for you.
The brain you pick for each agent decides three things: what it costs, how good it is, and where your data goes.
Kolibri changes the third one for a lot of people.
It is an open model you run on your own hardware, under a licence that lets you use it commercially.
Aleph Alpha says it was built with the EU AI Act, the General-Purpose AI Code of Practice and GDPR in mind from the ground up.
The model card confirms the company is a signatory of the EU GPAI Code of Practice.
So if you run agents for European clients who will not let their files touch an American cloud, Kolibri gives you a serious option that stays inside your own walls.
It also gives you a genuinely German-first brain, because more than a fifth of its pre-training data is German and it uses a tokenizer built for German words.
Aleph Alpha trained Kolibri on roughly 20 trillion tokens in pre-training, with about 62.5% English, about 23.9% German and about 13.6% code.
They added 3.44 trillion tokens of mid-training and 201 billion tokens of long-context training on top.
That is where the "about 24 trillion tokens" figure in my video comes from, and the official total is about 23.6 trillion.
The model is a mixture of experts with 384 experts in each of its 50 layers, and it picks 6 of them plus 1 shared expert for every token.
It also uses a hybrid attention design, where four layers look at nearby text and the fifth looks across the whole document.
That is why very long inputs stay affordable.
This is the part every Agent OS builder needs to hear before they get excited.
Only 3.46 billion parameters work on each token, but all 78 billion still have to sit in memory.
A viewer under my video said a 78B model needs roughly 100 GB of memory to run properly, and that most laptops cannot do it.
That is a fair warning, because the official FP8 weights alone are about 78 GB before you add room for context.
Aleph Alpha lists two 80 GB A100s, two H100s, one H200, one B200 or one B300 as the minimum.
The community has published smaller MLX and GGUF builds, and their converters list about 41 GiB for 4-bit MLX and about 24 GiB for 2-bit MLX.
That means a 64 GB Mac for the 4-bit build and a 36 GB Mac for the 2-bit build, with a bigger quality hit the lower you go.
Another viewer asked about a 2 GB graphics card, and that will not work.
There are four ways people try to run Kolibri, and only three of them work today.
| Route | What it needs | Status on 7 October 2026 | Best for in an Agent OS |
|---|---|---|---|
| vLLM plus Aleph Alpha's plugin | About 78 GB of GPU memory, such as 2× A100 80 GB or 1× H200. | This is the official route, and it works today. | A shared GPU box that serves several agents. |
| Community MLX build | A Mac with 64 GB or more for 4-bit, or 36 GB or more for 2-bit. | It works today with the launcher bundled in the download, and it is unofficial. | A single big Mac Studio running one agent. |
| Community GGUF plus patched llama.cpp | Plenty of system RAM and your own llama.cpp build. | It works today if you compile the patch, and it is unofficial. | Tinkerers who already build llama.cpp. |
| LM Studio or Ollama | The same memory as the builds above. | Stock apps cannot load it yet, and support requests are still open. | Nothing yet, so check again after each update. |
In my video I said you can run it through LM Studio, and I need to correct that for 7 October 2026.
Kolibri uses a brand new architecture, and LM Studio and Ollama depend on llama.cpp or MLX support that has not landed yet.
There is an open feature request on llama.cpp, an open pull request on mlx-lm and an open pull request on Ollama.
So search for Kolibri inside LM Studio after each update, and use one of the working routes until then.
This is the route I would use on any GPU box.
pip install 'aleph-alpha-inference>=1', which also installs the vLLM version it supports.vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 --reasoning-parser kolibri1 --tool-call-parser kolibri1 --enable-auto-tool-choice.http://localhost:8000/v1, which is a standard OpenAI-compatible address.Aleph Alpha also publishes a container image if you would rather run it in Docker.
This is the route for a Mac Studio or a high-memory MacBook Pro.
The standard mlx-lm library does not support Kolibri yet, so the community builds ship their own model file and a small launcher.
The launcher can start an OpenAI-compatible server too, so your agents talk to it the same way.
One converter reports around 52 to 56 tokens per second on an M1 Max, and I have not measured that myself.
This is the pairing I talked about in the video, where Kolibri becomes a free local brain for Hermes Agent.
The model card says the official server uses Hermes-style tool calling, which makes the fit even neater.
hermes model inside that profile and point it at your local endpoint.Watch out for the reasoning default on some community builds, because their converters say it falls back to high effort, which feels slow.
I do not run one model for everything, and you should not either.
| Agent OS lane | Good model choice | Why |
|---|---|---|
| German client work and long documents | Kolibri on a big Mac or GPU box. | It has native German, a long context window and tool calling, and the data stays in-house. |
| Fast everyday sub-agent jobs | LFM 2.5 2.6B. | It is the best local model I have tested with Hermes so far, and it is crazy fast. |
| Lightweight free local helper | Gemma 4. | It runs on far smaller machines, which is why I used it in my Run Hermes Free Forever video. |
| Hard planning and final checks | A frontier hosted model. | The heavy reasoning stays with the strongest brain, and the local lanes do the volume. |
I will be honest about my own setup, because it matters here.
I run a Mac Studio, and I do not actually run that much local AI.
The best local model I have tested with Hermes is still LFM 2.5 at 2.6 billion parameters, because it is crazy fast and it was trained with Hermes Agent in mind.
So Kolibri is not an automatic upgrade, and you should test it against your current local brain on your own jobs.
Every number below comes from Aleph Alpha's own model card, run on their own evaluation framework with Kolibri at high reasoning effort.
| Vendor-reported score | Kolibri | Qwen3.6 35B-A3B | Gemma 4 26B-A4B | Nemotron 3 Super | Mistral Small 4 | Qwen 3.8 27B (dense) |
|---|---|---|---|---|---|---|
| Overall, German | 70.8 | 67.3 | 66.3 | 67.9 | 61.4 | 79.9 |
| Overall, English | 75.5 | 71.4 | 71.9 | 73.0 | 63.1 | 80.2 |
| Agentic average | 63.4 | 62.1 | 54.6 | 54.9 | 40.7 | 66.7 |
That supports what I said in the video.
Kolibri has the best German overall score among the mixture-of-experts models in the table.
It also beats Qwen 3.6, Nemotron 3 Super and Mistral Small 4 on the multi-step tool-calling average, and Mistral Small 4 comes last of that group.
A tweet I showed claimed it beats Qwen, Nemotron and Mistral in maths, science and coding, and the vendor table broadly backs that up against those models.
But the dense Qwen 3.8 27B beats Kolibri on Aleph Alpha's own overall scores, which is exactly what one commenter told me.
Kolibri uses far fewer active parameters per token, so the real question is which one wins on your work at a cost you can live with.
Test it yourself, every time.
These limits decide where Kolibri belongs in your stack.
Add it if you run agents for German-speaking or European clients who care about where their data lives.
Add it if you already have a GPU server or a 64 GB-plus Mac sitting there.
Add it if your agents read long documents and you want to stop chopping them into pieces.
Skip it if you are on a 16 GB laptop, because a smaller local model will do more for you.
If you want the whole Agent OS with the local lanes already wired, it lives inside the AI Profit Boardroom.
My full free local Hermes setup is in the Hermes free-forever guide.
Kolibri AI is Kolibri 1, an open-weight mixture-of-experts model from Aleph Alpha in Germany. It has 78.1 billion total parameters, about 3.46 billion active per token, native German and English, reasoning and tool calling, and an Apache 2.0 licence. It was released on 3 October 2026.
The official spelling is Kolibri, the German word for hummingbird. Colibri AI and Calibri AI are common alternative spellings for the same model.
The official FP8 weights are about 78 GB, and Aleph Alpha lists two 80 GB A100s or one H200 as the minimum. Community 4-bit MLX builds need a 64 GB Mac, and the smallest community builds need about 36 GB.
Not with the stock apps as of 7 October 2026, because support for Kolibri's new architecture has not landed in llama.cpp, mlx-lm or Ollama yet. Use vLLM with Aleph Alpha's plugin or a community MLX build for now.
Yes. It supports Hermes-style tool calling and runs behind an OpenAI-compatible server, so a Hermes profile can point at the local endpoint. Test it against a fast small model like LFM 2.5 before you switch.
Check your memory first, start with one agent on one small job, and keep a human approval step on anything that matters.
That is how you turn Kolibri AI into a private brain your Agent OS can actually rely on.
About Julian
I'm Julian Goldie, an AI entrepreneur, SEO expert and the founder of the AI Profit Boardroom, which has 3,400+ members.
I help business owners scale with AI agents, automation and SEO.
My YouTube channel has 400,000+ subscribers, and I run Goldie Agency, a seven-figure SEO agency.
→ Run Hermes Free Forever With A Local Model
→ Hermes Agent + Qwen 3.8 27B: The Free Local AI Brain
→ Hermes + Gemma 4: Agentic Workflows On A Free Local Model
The installable system behind every guide on this site — dashboard, agents, memory and pipelines — updated as new tools land, with 3,400+ members and four coaching calls a week.
Join the AI Profit Boardroom →📺 Video notes + links to the tools 👉
🎥 Learn how I make these videos 👉