Xiaomi just dropped MiMo-V2.6, and it is the strongest open model anyone has ever released.
Two brains landed at once: a trillion-parameter Pro that sits one point behind GPT-5.6 Sol, and a Flash that costs pennies and still writes working code.
I'm going to show you how to plug both into Hermes and the Agent OS in about ten minutes, so your agents get a frontier brain without the frontier bill.
You'll see it build a full 3D racing game, run a real task inside my dashboard, and answer in under three seconds.
There is one thing you have to get right when you pick between Pro and Flash, and I'll show you exactly where the line is.
By the end, you'll have two new agents running that most people don't even know exist yet.
The big brain wakes only the parts it needs. The small one is built for speed. Both feed the same winged messenger, which is Hermes. That's the whole guide in one picture.
MiMo-V2.6-Pro scored 46 on the Artificial Analysis Intelligence Index, the highest any open-weights model has ever reached.
You can download it, and you can run it for less than a dollar per million words out.
Index scores as reported by Artificial Analysis on launch day. Closed models on the left, the open one on the right.
What you're watching: me driving the game Pro wrote while I was writing this section. Real keypresses, no fixes from me. That run cost 1.7 cents.
Pro is over a trillion parameters but only wakes about 42 billion per word. Flash is 309 billion and wakes 15 billion.
Both read text, images, video and audio, and both hold a million tokens at once.
Parameter counts from Xiaomi's model card and OpenRouter. "Active" is how much of the brain lights up for each word, and that is what you actually pay for.
What you're watching: the cheap model's game after it repaired itself. Asteroids in, shots out, score climbing. A step down from Pro's polish, which is exactly the trade you're making.
These are Xiaomi's own launch numbers, so hold them loosely. I'll show you where it loses further down.
Bars grow to the real scores as you scroll. Green is MiMo-V2.6-Pro, pink is the best closed rival on that test.
Xiaomi held prices flat from the last generation while the score jumped 20 points, and repeat prompts get a 99% cache discount.
I pulled these off the OpenRouter model list before writing this. My Pro game cost 1.7 cents, the Flash game 0.7 cents.
Most people run their agents on one expensive brain.
Every tiny task gets the frontier price, even "rename these files".
So they get scared of the meter, and the agents sit idle.
Or they switch to a cheap model, and the hard jobs quietly come back wrong.
Either way you're paying twice: once in money, once in work you have to redo.
Two open brains at two prices, in one operating system, break that cycle for good.
You don't swap anything. A Hermes profile is one folder with one line that names the model.
Add it, and the Agent OS lists it in the chat picker on its own. The rest of your setup never moves.
Pro has 256 small expert brains inside it, and each word you send wakes only eight of them.
That is why a model this big can cost less than a dollar per million words.
Mixture of experts, drawn honestly: 256 experts, 8 awake per word for Pro. Flash uses the same trick with a smaller crew. Numbers from Xiaomi's model card.
Here is the literal path, with the real timings from my machine.
~/.hermes/profiles/mimo-flash/ holds one config file that says which model to call and where.-p mimo-flash flag, so the right brain answers, not the default one.So when someone asks "but what IS it?": it's two more employees in Hermes, one senior and one fast, hired with one folder each.
Clone your working profile so the keys come along, then point each copy at its MiMo model.
What you're watching: my real session, replayed at reading pace. Create, edit the model line, check status.
Then open each profile's config.yaml and set the model block like this (swap pro for flash in the second one):
You don't need one. Both models are on OpenRouter, so the key already in your Hermes profile reaches them.
If you'd rather go direct, Xiaomi's own platform works too, but nothing in this guide needs it.
Hermes puts the model and provider in the system prompt, so a one-line question proves the routing.
What you're watching: both profiles answering with their real model id, then the Agent OS profile list showing both. Real outputs from my session.
Checkpoint: you should see xiaomi/mimo-v2.6-flash in the answer. If it names a different model, the config line didn't save.
I asked it to read a facts file and write a full markdown brief, table and all. It did it in under nine seconds.
What you're watching: the real Agent OS Hermes tab. The profile picker shows mimo-flash, the task goes in, Hermes thinks, and the reply confirms the file it wrote: 23 lines.
Checkpoint: open the file it wrote. Mine had the title, the table and both bullet lists, with nothing invented beyond the facts I gave it.
You can follow the three steps above yourself. Or get the whole operating system done, with Hermes, Claude, OpenClaw and Free Claude Code in one dashboard and every new model added the week it ships.
No. The everyday 90% runs on a free local model on your own machine, and free APIs slot in as profiles for more.
For frontier work it drives the CLIs you already pay for, so you're not paying twice, and this guide just added two brains that cost cents. Inside the Boardroom there are full token-efficiency tutorials so you never think about the meter again.
On a few tests the gap to the closed models is real, and you should know them before you route serious work.
Both models think before they answer. My Pro game used 4,986 reasoning tokens on top of 14,000 of code, and Flash used 12,952. Give them a big output budget, or the answer gets cut off mid-file.
The index score is independent. The benchmark table is Xiaomi's. Wait a week for the outside testers before you bet a client project on a single number.
Flash's game ran first time but the camera framing was off. One round of feedback and it fixed itself, but that is the trade: a third of the price, a step down in first-shot polish. Use it where "works" matters more than "beautiful".
Start every job on Flash. If it comes back wrong twice, or the job touches money, clients or security, send it to Pro.
The same rule I run for DeepSeek Flash and Pro. It works here because both MiMo brains share the same tools and the same 1M-token memory.
Xiaomi's MiMo team, led by Luo Fuli, spent six months scaling reinforcement learning after V2.5, and put the training dashboard on a public stream four days before launch.
"spent nearly six months studying how far reinforcement learning could scale after the release of MiMo-V2.5."
— Luo Fuli, head of Xiaomi MiMo, on the 18 September 2026 training livestream (via TechNode)
The weights, the technical report, the RL environments and the training code all shipped together. That is why the open-model crowd is treating this one differently.
Wrong: "Open models are always a year behind the closed ones."
Right: One point behind GPT-5.6 Sol on an independent index, and you can download it today.
Wrong: "Cheap means dumb. I'll get what I pay for."
Right: Flash wrote a working 3D game and a correct brief with a table for under a cent each. Cheap now means "wakes fewer experts", not "knows less".
Wrong: "Adding a model means rebuilding my whole agent setup."
Right: Two commands and one edited line. The Agent OS picked both profiles up without a restart.
Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.
Read the 158-page wins doc →Flash for the everyday work, Pro for the hard calls, both under a dollar and both open.
The people who wire new models in the week they land are the ones who'll be way ahead when all of this settles. Every profile you add compounds.
Your moveMiMo-V2.6 gives you a frontier-class brain for pennies. The Boardroom gives you the year I spent building everything around it: the operating system, the routing rules, the memory, and the people who fix your setup live on the calls.
Readers bookmark this page and keep paying one bill for everything. Operators join, install the Agent OS this week, and have Flash and Pro running client work by Friday.
Create the Flash profile first. Decide about Pro after it's answered you. I'll see you in the next one.