Run all of this inside the full Agent OS — join the AI Profit Boardroom →

agentos.guide › blog

What Is Kolibri AI? A Quick-Start For Running It Inside Your Agent OS

By Julian Goldie · 7 October 2026 · agentos.guide

A glowing bronze hummingbird hovering over a dark server rack with German and English text streams flowing between agent nodes

Kolibri AI is a free, open-weight model from Aleph Alpha in Germany, with 78.1 billion parameters, about 3.46 billion active per token, and native German and English.

People also search for it as Colibri, but the official name is Kolibri, which means hummingbird in German.

The question I get most is whether it can be a free, private brain for your agents.

It can, if you have the memory for it, so here is the quick-start, the honest hardware bill and exactly where it fits in an Agent OS.

What Kolibri AI is, in one minute

Kolibri AI is the open-weight model Aleph Alpha released on 3 October 2026, and its official name is Kolibri 1.

Kolibri is the German word for a hummingbird, which fits a big model that only flaps a small part of itself for each word.

You will also see it written as Colibri, and the auto-captions on my own video even spelled it Calibri.

Here is the spec sheet straight from the official Hugging Face model card.

SpecKolibri 1 (official model card)
MakerAleph Alpha, a German AI company based in Heidelberg.
Hugging Face repoAleph-Alpha/Kolibri-1 holds the FP8 weights, and Aleph-Alpha/Kolibri-1-BF16 holds full precision.
ReleasedIt was released on 3 October 2026.
LicenceIt uses Apache 2.0, and the repo is not gated.
SizeIt has 78.1 billion total parameters, with about 3.46 billion active per token.
LanguagesIt speaks German and English natively.
ContextIt reads 262,144 tokens natively, and Aleph Alpha validated it up to 1,048,576.
Knowledge cutoffIts built-in knowledge stops at 18 June 2026.
Agent featuresIt has a reasoning mode with effort levels and Hermes-style tool calling.
Input and outputIt is text only.

In my video I rounded the active parameters to "3 billion", and the exact official figure is 3.46 billion.

I also said it handles up to 1 million tokens, which is right, but Aleph Alpha recommends staying at or under 262,144 tokens for speed and for complex tasks.

🔥 Want the exact local models in the Agent OS setup I run inside my Agent OS?

Inside the AI Profit Boardroom you get the installable Agent OS, step-by-step video tutorials, four coaching calls a week and 3,400+ members building real automations.

→ Get access here

Why Kolibri AI matters for an Agent OS

An Agent OS is just a set of agents, tools and memory that run your work for you.

The brain you pick for each agent decides three things: what it costs, how good it is, and where your data goes.

Kolibri changes the third one for a lot of people.

It is an open model you run on your own hardware, under a licence that lets you use it commercially.

Aleph Alpha says it was built with the EU AI Act, the General-Purpose AI Code of Practice and GDPR in mind from the ground up.

The model card confirms the company is a signatory of the EU GPAI Code of Practice.

So if you run agents for European clients who will not let their files touch an American cloud, Kolibri gives you a serious option that stays inside your own walls.

It also gives you a genuinely German-first brain, because more than a fifth of its pre-training data is German and it uses a tokenizer built for German words.

How Kolibri AI was built, in plain English

Aleph Alpha trained Kolibri on roughly 20 trillion tokens in pre-training, with about 62.5% English, about 23.9% German and about 13.6% code.

They added 3.44 trillion tokens of mid-training and 201 billion tokens of long-context training on top.

That is where the "about 24 trillion tokens" figure in my video comes from, and the official total is about 23.6 trillion.

The model is a mixture of experts with 384 experts in each of its 50 layers, and it picks 6 of them plus 1 shared expert for every token.

It also uses a hybrid attention design, where four layers look at nearby text and the fifth looks across the whole document.

That is why very long inputs stay affordable.

The memory reality check before you plan anything

This is the part every Agent OS builder needs to hear before they get excited.

Only 3.46 billion parameters work on each token, but all 78 billion still have to sit in memory.

A viewer under my video said a 78B model needs roughly 100 GB of memory to run properly, and that most laptops cannot do it.

That is a fair warning, because the official FP8 weights alone are about 78 GB before you add room for context.

Aleph Alpha lists two 80 GB A100s, two H100s, one H200, one B200 or one B300 as the minimum.

The community has published smaller MLX and GGUF builds, and their converters list about 41 GiB for 4-bit MLX and about 24 GiB for 2-bit MLX.

That means a 64 GB Mac for the 4-bit build and a 36 GB Mac for the 2-bit build, with a bigger quality hit the lower you go.

Another viewer asked about a 2 GB graphics card, and that will not work.

Kolibri AI quick-start: the four routes

There are four ways people try to run Kolibri, and only three of them work today.

RouteWhat it needsStatus on 7 October 2026Best for in an Agent OS
vLLM plus Aleph Alpha's pluginAbout 78 GB of GPU memory, such as 2× A100 80 GB or 1× H200.This is the official route, and it works today.A shared GPU box that serves several agents.
Community MLX buildA Mac with 64 GB or more for 4-bit, or 36 GB or more for 2-bit.It works today with the launcher bundled in the download, and it is unofficial.A single big Mac Studio running one agent.
Community GGUF plus patched llama.cppPlenty of system RAM and your own llama.cpp build.It works today if you compile the patch, and it is unofficial.Tinkerers who already build llama.cpp.
LM Studio or OllamaThe same memory as the builds above.Stock apps cannot load it yet, and support requests are still open.Nothing yet, so check again after each update.

In my video I said you can run it through LM Studio, and I need to correct that for 7 October 2026.

Kolibri uses a brand new architecture, and LM Studio and Ollama depend on llama.cpp or MLX support that has not landed yet.

There is an open feature request on llama.cpp, an open pull request on mlx-lm and an open pull request on Ollama.

So search for Kolibri inside LM Studio after each update, and use one of the working routes until then.

Route 1: the official vLLM server

This is the route I would use on any GPU box.

Aleph Alpha also publishes a container image if you would rather run it in Docker.

Route 2: a community MLX build on a big Mac

This is the route for a Mac Studio or a high-memory MacBook Pro.

The standard mlx-lm library does not support Kolibri yet, so the community builds ship their own model file and a small launcher.

The launcher can start an OpenAI-compatible server too, so your agents talk to it the same way.

One converter reports around 52 to 56 tokens per second on an M1 Max, and I have not measured that myself.

How to plug Kolibri AI into Hermes inside your Agent OS

This is the pairing I talked about in the video, where Kolibri becomes a free local brain for Hermes Agent.

The model card says the official server uses Hermes-style tool calling, which makes the fit even neater.

Watch out for the reasoning default on some community builds, because their converters say it falls back to high effort, which feels slow.

Where I would put Kolibri in the lanes

I do not run one model for everything, and you should not either.

Agent OS laneGood model choiceWhy
German client work and long documentsKolibri on a big Mac or GPU box.It has native German, a long context window and tool calling, and the data stays in-house.
Fast everyday sub-agent jobsLFM 2.5 2.6B.It is the best local model I have tested with Hermes so far, and it is crazy fast.
Lightweight free local helperGemma 4.It runs on far smaller machines, which is why I used it in my Run Hermes Free Forever video.
Hard planning and final checksA frontier hosted model.The heavy reasoning stays with the strongest brain, and the local lanes do the volume.

I will be honest about my own setup, because it matters here.

I run a Mac Studio, and I do not actually run that much local AI.

The best local model I have tested with Hermes is still LFM 2.5 at 2.6 billion parameters, because it is crazy fast and it was trained with Hermes Agent in mind.

So Kolibri is not an automatic upgrade, and you should test it against your current local brain on your own jobs.

Kolibri AI benchmarks, flagged as vendor-reported

Every number below comes from Aleph Alpha's own model card, run on their own evaluation framework with Kolibri at high reasoning effort.

Vendor-reported scoreKolibriQwen3.6 35B-A3BGemma 4 26B-A4BNemotron 3 SuperMistral Small 4Qwen 3.8 27B (dense)
Overall, German70.867.366.367.961.479.9
Overall, English75.571.471.973.063.180.2
Agentic average63.462.154.654.940.766.7

That supports what I said in the video.

Kolibri has the best German overall score among the mixture-of-experts models in the table.

It also beats Qwen 3.6, Nemotron 3 Super and Mistral Small 4 on the multi-step tool-calling average, and Mistral Small 4 comes last of that group.

A tweet I showed claimed it beats Qwen, Nemotron and Mistral in maths, science and coding, and the vendor table broadly backs that up against those models.

But the dense Qwen 3.8 27B beats Kolibri on Aleph Alpha's own overall scores, which is exactly what one commenter told me.

Kolibri uses far fewer active parameters per token, so the real question is which one wins on your work at a cost you can live with.

Test it yourself, every time.

Kolibri AI limits that change how you design agents

These limits decide where Kolibri belongs in your stack.

Who should add Kolibri AI to their Agent OS

Add it if you run agents for German-speaking or European clients who care about where their data lives.

Add it if you already have a GPU server or a 64 GB-plus Mac sitting there.

Add it if your agents read long documents and you want to stop chopping them into pieces.

Skip it if you are on a 16 GB laptop, because a smaller local model will do more for you.

If you want the whole Agent OS with the local lanes already wired, it lives inside the AI Profit Boardroom.

My full free local Hermes setup is in the Hermes free-forever guide.

Also On Our Network

FAQ

What is Kolibri AI?

Kolibri AI is Kolibri 1, an open-weight mixture-of-experts model from Aleph Alpha in Germany. It has 78.1 billion total parameters, about 3.46 billion active per token, native German and English, reasoning and tool calling, and an Apache 2.0 licence. It was released on 3 October 2026.

Is it Kolibri or Colibri?

The official spelling is Kolibri, the German word for hummingbird. Colibri AI and Calibri AI are common alternative spellings for the same model.

How much memory does Kolibri AI need?

The official FP8 weights are about 78 GB, and Aleph Alpha lists two 80 GB A100s or one H200 as the minimum. Community 4-bit MLX builds need a 64 GB Mac, and the smallest community builds need about 36 GB.

Can I run Kolibri AI in LM Studio or Ollama?

Not with the stock apps as of 7 October 2026, because support for Kolibri's new architecture has not landed in llama.cpp, mlx-lm or Ollama yet. Use vLLM with Aleph Alpha's plugin or a community MLX build for now.

Can Kolibri AI run Hermes Agent?

Yes. It supports Hermes-style tool calling and runs behind an OpenAI-compatible server, so a Hermes profile can point at the local endpoint. Test it against a fast small model like LFM 2.5 before you switch.

Final word on Kolibri AI

Check your memory first, start with one agent on one small job, and keep a human approval step on anything that matters.

That is how you turn Kolibri AI into a private brain your Agent OS can actually rely on.

About Julian

I'm Julian Goldie, an AI entrepreneur, SEO expert and the founder of the AI Profit Boardroom, which has 3,400+ members.

I help business owners scale with AI agents, automation and SEO.

My YouTube channel has 400,000+ subscribers, and I run Goldie Agency, a seven-figure SEO agency.

→ Get my best AI training inside the AI Profit Boardroom

Related guides on agentos.guide

→ Run Hermes Free Forever With A Local Model

→ Hermes Agent + Qwen 3.8 27B: The Free Local AI Brain

→ Hermes + Gemma 4: Agentic Workflows On A Free Local Model

Get the whole Agent OS

The installable system behind every guide on this site — dashboard, agents, memory and pipelines — updated as new tools land, with 3,400+ members and four coaching calls a week.

Join the AI Profit Boardroom →

📺 Video notes + links to the tools 👉

🎥 Learn how I make these videos 👉

🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉

← all guides on agentos.guide