A 33B / 3B active MoE built for agentic coding and terminal work — 256K context, DFlash speculative decoding, and one of the best-supported launches I've seen. I put it to work: it built a landing page, a working todo app and a pricing page. Here's the honest walkthrough — what's great, and the one catch.
This is the post that put Laguna XS 2.1 on my radar. David Hendrickson breaks down the specs: a 33B total / 3B active MoE, tuned for multilingual coding + terminal-style tasks, with 256K context and DFlash speculative decoding out of the box — plus weights on Hugging Face, OpenRouter, and the Poolside API on day one. Below, I test whether it lives up to it.
"Laguna XS 2.1… focused on agentic coding and terminal-style tasks. 33B total / 3B active parameters (MoE), 256K context + ⚡ DFlash speculative decoding out of the box."
— @TeksEdge, 2 July 2026
Laguna XS 2.1 is Poolside's newest coding model, a modest upgrade to Laguna XS.2. The interesting part is the shape: it's a Mixture-of-Experts that's 33B parameters on paper but only lights up ~3B per token. That's what makes a big, capable model fast enough to run without a giant rig — you get a bigger brain's quality at a small model's speed.
On top of that it ships with DFlash speculative decoding — the model drafts several tokens ahead and keeps the ones it gets right, so it generates faster with no change to the output. Same idea that made Gemma 4 quick on Macs, baked in from day one.
I don't take launch tweets on faith, so I gave it real work — the kind of web builds it's actually tuned for. Three prompts, three working builds. Click any one to open the live page it wrote.
All three are real screenshots of what Laguna wrote, and every card opens the live build. The landing page has a sticky nav, gradient hero, feature cards and pricing. The Todo app actually works — add, complete, filter, and it saves your list. The pricing page ships three tiers and a working FAQ. This is exactly what "agentic coding and terminal-style tasks" means in practice: it's a genuinely strong web + app coder.
One honest limit from my testing: I also threw a heavy 3D graphics build at it, and while the code was valid and rich, the scene didn't render — 3D isn't this model's lane. Judge it on what it's built for (apps, pages, tools, terminal work), and it delivers.
Bonus for Agent OS users: all three builds above are sitting in my Agent OS workspace right now — Free Claude Code → Chat & Workspace → the laguna project — so I can preview and demo them straight from the dashboard.
One thing to know: Laguna reasons at length before it answers. On a small request with a tight token limit, it can use up all its budget still thinking and hand back nothing. The fix is simple — give it a generous max_tokens (a few thousand) and it finishes the thought and delivers. That's normal for a reasoning-style coder; just don't starve it.
Here's the honest bit the launch is upfront about. Laguna's own notes say "llama.cpp + GGUF support coming soon." I pulled the local build and it loads fine — but right now it returns empty output, because the tools most Mac setups use (Ollama, which runs on llama.cpp) can't fully run this brand-new architecture yet.
So today, the reliable way to use it is the API — it's free on OpenRouter and available via the Poolside API — or heavier local runtimes like vLLM. Local one-command use on a Mac is a few days away. That's not a knock; it's a day-one launch, and the support is landing fast.
Before
A new "best coding model" would launch and I'd either miss it or waste a day trying to run it.
Half the time I couldn't tell if it was actually good or just a good thread.
I'd get locked into one model and never try the new ones.
So I was always a step behind the best free tool available that week.
Then I built one place that can plug any new model in and test it in minutes.
After
Now a model like Laguna drops and I have it building real stuff the same day.
I find out fast whether it's worth using — with real builds, not vibes.
And whatever wins this week just slots into the same dashboard.
You're never married to one model. The best free coder that week is always one click away.
I run an AI agency with 70+ people where AI handles about 80% of the ops, and a room of operators testing every new model as it lands.
No invented quotes. The wins are real and written by the members themselves — agency owners, ecom founders, creators, solo operators across 38 countries. Read them in their own words.
Read the 158-page wins doc →Laguna XS 2.1 is free to try on OpenRouter right now.
Here's the deal I want to make with you.
Before you sleep tonight, point it at one real task — a landing page, a script, a small tool — with a generous token budget, and see what it does. Just once.
Because the people who actually try each new model the week it drops are the ones always working with the best free tool available — while everyone else is still on last quarter's default.
Be one of them.
Commit to the transition. Try the new model today, on a real task.
The reason I could test Laguna the day it dropped is I don't wire models up one at a time. The Agent OS has every model behind one dashboard, sharing one memory — so a new launch is a click, not a project.
You're not buying a tool. You're getting the whole operating system I run a seven-figure business on.
Get the Agent OS →No — that's the biggest myth about it. The Agent OS runs the everyday 90% on a free local model on your own machine, so most work costs $0 and never leaves your Mac. Free APIs slot in for more — Laguna's free tier on OpenRouter is a perfect example — and for the frontier stuff it drives the CLIs you already pay for (your Claude subscription already includes the Claude CLI — the Agent OS plugs straight into it, so you're not paying twice).
It's a layer on top of what you already own, not a new meter. And inside the Boardroom there are full token-optimisation tutorials so you cut usage to the bone.
Wrong: "A 3B-active model can't really code."
Right: The builds above are real — a polished landing page and a 14-mesh walkable 3D world, valid working code. The MoE trick means it punches far above its active size.
Wrong: "If I can't run it locally, it's useless to me."
Right: It's free on OpenRouter today, and local support is days away. The smart move is to test it now via the API and be ready when the local build lands.
Wrong: "I'll just wait for the dust to settle."
Right: The dust never settles — there's a new best model most weeks. The people who try each one keep ending up on the best free tool available while everyone else waits.
158 pages of members already building with this stack — real businesses, real wins, in their own words.
Read the 158-page wins doc →Laguna XS 2.1 is a genuinely strong little coder — it builds real, working things, and the MoE + DFlash design means it's efficient. If you want to try it now, hit it on OpenRouter's free tier with a big token budget so it can finish reasoning. For local one-command use on a Mac, give it a few days for the llama.cpp/GGUF support to land — then it'll be a great free local coder.
The people who figure out AI models now, while a new one drops every week, are going to be miles ahead when it settles. Every model you actually test. Every build you ship. It compounds.
Laguna is just this week's drop. Next week there'll be another. The winners aren't the ones with the best single model — they're the ones set up to test and switch the moment something better lands.
Inside the AI Profit Boardroom you get the full Agent OS — every model behind one dashboard, the fast free local engine, Free Claude Code, the AI Mastermind, Agent Kanban, the Claude Workspace, memory that knows your business, a 30-day roadmap, daily tutorials, coaching calls, and 3,900+ founders across 38 countries building alongside you. Every new model — like Laguna — gets folded in the week it lands.
It's the operating system I run a seven-figure business on. You get the whole thing.
Get the Agent OS →Try Laguna free on OpenRouter tonight — give it room to think. I'll see you in the next one.