Sakana Fugu 1.1 just dropped — and it's unlike any AI model you've used.
You give it one prompt, and it sends a whole team of AI experts to build it for you.
I put that team through the hardest game builds on my benchmark.
It matched the best model on the planet on its very first try.
Every build is on this page — you can play them right now.
And later I'll show you the one feedback trick that turned its worst build into one of its best.
By the end, you'll know exactly how to put this expert team to work in your own stack.

Every model you use has the same hidden limit.
One brain does the whole job.
The same brain plans your game's physics, paints its graphics, and wires its scoreboard.
When the job gets big, that one brain starts dropping things.
The graphics land but the controls break. The physics work but the HUD never renders.
You've seen it — every long build that came back 90% right and 10% broken.
Sakana's answer: stop using one brain. Route the job to a team of experts and merge what comes back.
That's Fugu 1.1 — and today I put the team on real work.
It IS slower — my first build took 25 minutes of orchestrated thinking. The honest question is whether the output justifies it.
That's what the embedded builds below are for. Play them and judge the trade yourself.
You send Fugu one prompt, like any model.
Behind the API, an orchestrator reads the job and decides how many experts it needs — one for a simple task, up to three for a hard one.
Each expert works its slice. The orchestrator merges the results into one answer.
You never see the team. You just get one response — and a usage line showing the orchestration tokens the team burned.
Three practical facts before you touch it:
Effort goes up to "max". Levels are high / xhigh / max (max is new in 1.1, default xhigh). Higher effort = bigger expert teams, longer thinking, better hard-task output.
Temperature is ignored. The orchestrator runs its own show. Don't tune what isn't listening.
max_output_tokens caps only the FINAL answer. The experts' own thinking bills separately as orchestration tokens in the usage block — budget for the team, not just the answer.
Sakana's own launch numbers: 73.7 on SWE-Bench Pro and 93.2 on LiveCodeBench — vendor-reported, so treat them as claims. My bench below is what I can verify.
The key. One API key from console.sakana.ai, stored at ~/.config/sakana/fugu.key. Base URL https://api.sakana.ai/v1 — OpenAI-style chat completions work, and so does Anthropic's messages format.
One honest gotcha: the region wall. Sakana currently blocks the EU, UK and Switzerland at the edge — a bare 403 before your key is even checked. From an allowed region (Asia, US), the same request just works. If you're seeing 403 on everything, it's geography, not your key.
The Hermes profile. One profile in the Agent OS points at the Sakana endpoint with the key — then sakana-fugu -z "the job" --yolo dispatches Fugu like any other staff member, with tools.
The verification test. Never trust a wiring until an agent proves it end-to-end. I gave it a job with a checkable end state — write a file, verify it exists, report its exact size. Here's the run, unedited:
$ sakana-fugu -z "Create fugu-agent-demo.md … verify it exists on disk, report exact size." --yolo /Users/juliangoldie/.hermes/profiles/sakana-fugu/workspace/fugu-agent-demo.md Word count: 200 Exact size: 1,454 bytes $ ls -la …/workspace/fugu-agent-demo.md -rw------- 1 juliangoldie staff 1454 fugu-agent-demo.md ← matches. verified.
What you're looking at: the whole integration in one receipt. Fugu wrote the file, checked the disk itself, reported 1,454 bytes — and the filesystem agrees to the byte. That's what "working inside the Agent OS" means.
The whole wiring is one config file pointing at one URL with one key — and inside the Boardroom you get it pre-built.
If you can paste a key into a text file, you can run this.
Fugu 1.1 is running my standard GoldieBench gauntlet today: the same one-shot, single-file build prompts every frontier model has faced, scored 0–10 by the same judge.
The run auto-publishes — every build lands on the bench the moment it's judged. First result:
Dragon Realm — 8.6/10. The hardest open-world task on the bench. Fugu's orchestrator thought for 25 minutes and returned a 63,184-byte game: third-person swordfighter, six roaming ice-spike enemies, health AND stamina systems, a compass, a rune-collection objective, and a full control bar. Zero console errors. Play it:
What you're looking at: the actual file Fugu returned, served from the live bench — not a mockup. For scale: Claude Fable 5 scored 7.8 on this same brief, GPT-5.6 Sol 8.6. The rookie matched Sol's score on attempt one.
Then the other six priority games ran the same gauntlet. The full first-pass board:
Dragon Realm — 8.6. Open-world swordfighter: six enemies, stamina, compass, runes. Matched GPT-5.6 Sol.
Parachute — 8.6. A combat parachute drop the judge tagged as a task standout.
Flight sim — 8.4. Its fastest build of the set — 11.5 minutes.
Crypt — 7.8. Torch-lit dungeon crawler, solid but mid-field.
GTA-drive — 7.4 · GTA-foot — 6.8. The open-city briefs exposed it: playable but generic against the field's best.
Doom — 6.0 first pass → 8.4 after the fix loop. This is the orchestrator's party trick: my QA loop fed Fugu its own game's failures, and two fix rounds later the thin maze became a real 3D demon shooter. The weak spots aren't dead ends — they're one feedback pass from strong.
The honest read: 7.66 first-pass, and it climbs fast when you feed it feedback — the Doom jump alone shows what the expert team does with a concrete fix list. For a brand-new lab's flagship on the hardest briefs in the set, that's a serious debut — v1.0 averaged 7.94 across the FULL bench including easier tasks, and 1.1's full-board number is filling in overnight on the live scorecard.
What you're looking at: Fugu's second 8.6 — a full combat parachute drop, built in 13 minutes of orchestrated thinking. Both embeds on this page are the actual judged files, served from the live bench.
Every build Fugu 1.1 has completed on the bench so far, best first — real judge scores, real files, each one click from playing full screen. This gallery grows as the remaining tasks publish.
The always-current version lives on the live scorecard — this page refreshes with each batch.
Correct — which is why the pipeline doesn't stop at one. Seven priority games, a fix loop, then the whole 50-task bench, all landing on the public scorecard with playable files.
This guide links the live board precisely so you never have to take a static screenshot's word for it.
Everything Fugu produces today flows into one place: the Sakana Fugu project in the Agent OS workspace.
The verified agent demo file, every bench build as it's judged, and the sync script that keeps it current — all visible in Claude → Workspace → sakana-fugu, previewable in one click.
That's the standing rule for every model in the stack: no output vanishes into a chat log. It lands as a file, in the workspace, or it didn't happen.
No — that's the biggest myth about it. The everyday 90% runs on free local models on your own machine, and free APIs slot in for more.
Frontier work drives the subscriptions and keys you already own — Fugu here runs on the same Sakana key you'd use anyway; the Agent OS is the layer that turns it into an agent with tools and a workspace, not a second meter.
And inside the AI Profit Boardroom there are full token-optimisation tutorials — which matter double with an orchestrator that bills its expert team separately.
Wrong: "I should pick one model and stick with it."
Right: Every model on my bench wins somewhere and loses somewhere. The operators winning right now run a stack of switches and route each job to whichever brain — or team of brains — fits it.
Wrong: "Slow models aren't worth it."
Right: Fugu thought for 25 minutes and matched GPT-5.6 Sol's score on the hardest open-world task, first try. Slow is fine when it's unattended — that's what agents are for.
Wrong: "Vendor benchmarks tell me what I need to know."
Right: Sakana reports 73.7 SWE-Bench Pro. Useful signal — but the playable builds on the live bench are evidence you can drive. Always prefer receipts you can touch.
Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.
Read the 158-page wins doc →The bench is public, the builds are playable, and the workflow runs inside the same Agent OS 4,000+ founders use daily.
Get a key at console.sakana.ai (or use the OpenRouter route if the console's not available to you).
Check your region first. One curl to the chat endpoint. A 403 on everything = the EU/UK wall, not a broken key.
Store the key in a file with tight permissions — never in a script, never in a repo.
Create the agent profile — one config pointing at the endpoint, key from the environment.
Run the verification test. A file-write job with a checkable end state. No agent joins the stack without passing it.
Give it a real job at default effort — xhigh. Only reach for max when a hard task underdelivers.
Watch the usage block. Orchestration tokens are the real bill. Note what a typical job costs before you scale it.
Route by lane. Fugu gets the big, hard, unattended builds. Your fast daily driver keeps the quick chats.
Then you flip a different switch. That's the whole point of running models as profiles in one stack — Fugu is a staff member, not a foundation. Nothing in your workflow depends on any single vendor staying generous.
Key, profile, verification test. Fugu answers from inside your stack.
Run your real jobs through it at xhigh. Note where the expert team beats your daily driver.
Give it the long builds overnight. Slow + good + unattended = free capacity.
Compare effort levels on your workload. Keep max for the jobs that earn it.
One prompt, an orchestrator, up to three experts, one merged answer — and a separate bill for the team.
A real agent run with a byte-exact receipt, wired into the Agent OS through one profile.
8.6 on the bench's hardest open-world task, first attempt — embedded above, judged like every other model.
The EU/UK region wall, ignored temperature, orchestration billing, and the 25-minute patience tax.
The rest of the 50-task run publishes itself — the link is the truth, not a screenshot.
Big unattended builds go to the team. Quick chats stay fast. Nobody gets a monopoly.