Tested today · 30 July 2026 · Sakana Fugu Ultra 1.1 × Agent OS

Sakana Fugu 1.1, actually tested.

Sakana Fugu 1.1 just dropped — and it's unlike any AI model you've used.

You give it one prompt, and it sends a whole team of AI experts to build it for you.

I put that team through the hardest game builds on my benchmark.

It matched the best model on the planet on its very first try.

Every build is on this page — you can play them right now.

And later I'll show you the one feedback trick that turned its worst build into one of its best.

By the end, you'll know exactly how to put this expert team to work in your own stack.

A giant clockwork brass pufferfish floating like an airship above a moonlit Japanese bay, three glowing workshop chambers visible in its belly with robed artisans fully clothed at their benches, torii gate and paper lanterns below
THE PUFFERFISH — one prompt in, a team comes out Your promptone job The orchestratorroutes 1–3 experts Expert · gameplay Expert · graphics Expert · systems One buildmerged + shipped the expert-team fee is real — orchestration tokens bill separately — but so is the output quality below
why Sakana named it after the pufferfish — one fish, and it inflates into a team when threatened by a hard prompt
7.66
average across the 7 priority game tasks
1–3
expert agents routed per job
1M
token context window
2×8.6
top scores — open world + parachute combat
✦
I · the problem

The One-Brain Ceiling Problem.

Every model you use has the same hidden limit.

One brain does the whole job.

The same brain plans your game's physics, paints its graphics, and wires its scoreboard.

When the job gets big, that one brain starts dropping things.

The graphics land but the controls break. The physics work but the HUD never renders.

You've seen it — every long build that came back 90% right and 10% broken.

Sakana's answer: stop using one brain. Route the job to a team of experts and merge what comes back.

That's Fugu 1.1 — and today I put the team on real work.

Thinking it?"Multi-agent sounds like marketing for slow and expensive."

It IS slower — my first build took 25 minutes of orchestrated thinking. The honest question is whether the output justifies it.

That's what the embedded builds below are for. Play them and judge the trade yourself.

✦
II · how it works, in simple words

One fish that inflates into a team.

You send Fugu one prompt, like any model.

Behind the API, an orchestrator reads the job and decides how many experts it needs — one for a simple task, up to three for a hard one.

Each expert works its slice. The orchestrator merges the results into one answer.

You never see the team. You just get one response — and a usage line showing the orchestration tokens the team burned.

Three practical facts before you touch it:

1

Effort goes up to "max". Levels are high / xhigh / max (max is new in 1.1, default xhigh). Higher effort = bigger expert teams, longer thinking, better hard-task output.

2

Temperature is ignored. The orchestrator runs its own show. Don't tune what isn't listening.

3

max_output_tokens caps only the FINAL answer. The experts' own thinking bills separately as orchestration tokens in the usage block — budget for the team, not just the answer.

Sakana's own launch numbers: 73.7 on SWE-Bench Pro and 93.2 on LiveCodeBench — vendor-reported, so treat them as claims. My bench below is what I can verify.

THE EFFORT LADDER — how hard the team tries highfast · smaller team xhighthe defaultbalanced team + depth max — new in 1.1full expert teamlongest thinkingfor the hardest builds my 8.6 came at default xhigh — max is headroom this bench hasn’t even used yet
the ladder — climb only when a job underdelivers; every rung costs more orchestration tokens
✦
III · exactly how it works — wired into the agent os

From API key to working agent, step by step.

1

The key. One API key from console.sakana.ai, stored at ~/.config/sakana/fugu.key. Base URL https://api.sakana.ai/v1 — OpenAI-style chat completions work, and so does Anthropic's messages format.

2

One honest gotcha: the region wall. Sakana currently blocks the EU, UK and Switzerland at the edge — a bare 403 before your key is even checked. From an allowed region (Asia, US), the same request just works. If you're seeing 403 on everything, it's geography, not your key.

3

The Hermes profile. One profile in the Agent OS points at the Sakana endpoint with the key — then sakana-fugu -z "the job" --yolo dispatches Fugu like any other staff member, with tools.

4

The verification test. Never trust a wiring until an agent proves it end-to-end. I gave it a job with a checkable end state — write a file, verify it exists, report its exact size. Here's the run, unedited:

THE REGION WALL — same key, different door EU · UK · CHrequest sent… 403blocked at the edge Asia · USsame request… api.sakana.ai · 200Fugu answers
the gotcha that cost this test a week — 403 everywhere means your location, not your key
sakana-fugu · verified agent run · today ● unedited
$ sakana-fugu -z "Create fugu-agent-demo.md … verify it exists on disk, report exact size." --yolo

/Users/juliangoldie/.hermes/profiles/sakana-fugu/workspace/fugu-agent-demo.md
Word count: 200
Exact size: 1,454 bytes

$ ls -la …/workspace/fugu-agent-demo.md
-rw-------  1 juliangoldie  staff  1454  fugu-agent-demo.md   ← matches. verified.

What you're looking at: the whole integration in one receipt. Fugu wrote the file, checked the disk itself, reported 1,454 bytes — and the filesystem agrees to the byte. That's what "working inside the Agent OS" means.

Thinking it?"I'm not technical enough to wire an API into an agent stack."

The whole wiring is one config file pointing at one URL with one key — and inside the Boardroom you get it pre-built.

If you can paste a key into a text file, you can run this.

IV · the sources

The receipts, all clickable.

Official sources + live results ↓
One brain drops things. A team catches them — you just pay the team.
V · the bench test — live and still publishing

Real builds, scored by the same judge as every other model.

Fugu 1.1 is running my standard GoldieBench gauntlet today: the same one-shot, single-file build prompts every frontier model has faced, scored 0–10 by the same judge.

The run auto-publishes — every build lands on the bench the moment it's judged. First result:

Dragon Realm — 8.6/10. The hardest open-world task on the bench. Fugu's orchestrator thought for 25 minutes and returned a 63,184-byte game: third-person swordfighter, six roaming ice-spike enemies, health AND stamina systems, a compass, a rune-collection objective, and a full control bar. Zero console errors. Play it:

Fugu 1.1's first build · Dragon Realm · judged 8.6 · WASD to move, click to strike

⛶ open full screen

What you're looking at: the actual file Fugu returned, served from the live bench — not a mockup. For scale: Claude Fable 5 scored 7.8 on this same brief, GPT-5.6 Sol 8.6. The rookie matched Sol's score on attempt one.

Then the other six priority games ran the same gauntlet. The full first-pass board:

1

Dragon Realm — 8.6. Open-world swordfighter: six enemies, stamina, compass, runes. Matched GPT-5.6 Sol.

2

Parachute — 8.6. A combat parachute drop the judge tagged as a task standout.

3

Flight sim — 8.4. Its fastest build of the set — 11.5 minutes.

4

Crypt — 7.8. Torch-lit dungeon crawler, solid but mid-field.

5

GTA-drive — 7.4 · GTA-foot — 6.8. The open-city briefs exposed it: playable but generic against the field's best.

6

Doom — 6.0 first pass → 8.4 after the fix loop. This is the orchestrator's party trick: my QA loop fed Fugu its own game's failures, and two fix rounds later the thin maze became a real 3D demon shooter. The weak spots aren't dead ends — they're one feedback pass from strong.

The honest read: 7.66 first-pass, and it climbs fast when you feed it feedback — the Doom jump alone shows what the expert team does with a concrete fix list. For a brand-new lab's flagship on the hardest briefs in the set, that's a serious debut — v1.0 averaged 7.94 across the FULL bench including easier tasks, and 1.1's full-board number is filling in overnight on the live scorecard.

The parachute drop · judged 8.6 · click to play

⛶ open full screen

What you're looking at: Fugu's second 8.6 — a full combat parachute drop, built in 13 minutes of orchestrated thinking. Both embeds on this page are the actual judged files, served from the live bench.

V·b — every build so far

All 7 judged builds, playable.

Every build Fugu 1.1 has completed on the bench so far, best first — real judge scores, real files, each one click from playing full screen. This gallery grows as the remaining tasks publish.

Dragonrealm · 8.6 GAME⛶ full screen
Parachute · 8.6 GAME⛶ full screen
Flightsim · 8.4 GAME⛶ full screen
Doom · 8.4 GAME⛶ full screen
Crypt · 7.8 GAME⛶ full screen
Gtadrive · 7.4 GAME⛶ full screen
Gtafoot · 6.8 GAME⛶ full screen

The always-current version lives on the live scorecard — this page refreshes with each batch.

DRAGON REALM — same brief, same judge, first attempt Fugu Ultra 1.1 · 8.6 — first ever attempt GPT-5.6 Sol · 8.6 Claude Fable 5 · 7.8
one data point so far — honest label: the full run is still publishing; the live board is the source of truth
Thinking it?"One good score proves nothing."

Correct — which is why the pipeline doesn't stop at one. Seven priority games, a fix loop, then the whole 50-task bench, all landing on the public scorecard with playable files.

This guide links the live board precisely so you never have to take a static screenshot's word for it.

✦
VI · the agent os workspace

Every output lands where you can touch it.

Everything Fugu produces today flows into one place: the Sakana Fugu project in the Agent OS workspace.

The verified agent demo file, every bench build as it's judged, and the sync script that keeps it current — all visible in Claude → Workspace → sakana-fugu, previewable in one click.

That's the standing rule for every model in the stack: no output vanishes into a chat log. It lands as a file, in the workspace, or it didn't happen.

NOTHING VANISHES — every output lands as a file Fugu buildsgames · files · docs GoldieBench · liveauto-published + judged Agent OS workspace · sakana-fuguClaude → Workspace → one-click preview chat logs evaporate — files compound. the workspace is where the work stays yours
the standing rule for every model in the stack — output lands as a file or it didn’t happen
Thinking it?"Doesn't running the Agent OS burn a fortune in tokens?"

No — that's the biggest myth about it. The everyday 90% runs on free local models on your own machine, and free APIs slot in for more.

Frontier work drives the subscriptions and keys you already own — Fugu here runs on the same Sakana key you'd use anyway; the Agent OS is the layer that turns it into an agent with tools and a workspace, not a second meter.

And inside the AI Profit Boardroom there are full token-optimisation tutorials — which matter double with an orchestrator that bills its expert team separately.

Skip the setup

Get the Sakana Fugu 1.1 setup built for you.

You can wire this yourself with the steps above. Or get it done inside the Agent Operating System — the profile, the bench harness, and the workspace pattern pre-connected.

The full Agent OS zip — every frontier model wired as a switch
The Hermes profile pattern for any new API in two minutes
Coaching calls where we wire your keys and run your first agents together
A room of 4,000+ operators comparing real model results daily
The prompts, the SOPs, and a member map for your city
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Set up in an afternoon · used in 38 countries · new models added the week they ship
✦
VII · three beliefs to drop

What's actually holding you back.

Wrong: "I should pick one model and stick with it."

Right: Every model on my bench wins somewhere and loses somewhere. The operators winning right now run a stack of switches and route each job to whichever brain — or team of brains — fits it.

Wrong: "Slow models aren't worth it."

Right: Fugu thought for 25 minutes and matched GPT-5.6 Sol's score on the hardest open-world task, first try. Slow is fine when it's unattended — that's what agents are for.

Wrong: "Vendor benchmarks tell me what I need to know."

Right: Sakana reports 73.7 SWE-Bench Pro. Useful signal — but the playable builds on the live bench are evidence you can drive. Always prefer receipts you can touch.

Don't take my word for it

Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.

Read the 158-page wins doc →
✦
VIII · the receipts

Who's already running this.

The bench is public, the builds are playable, and the workflow runs inside the same Agent OS 4,000+ founders use daily.

4,000+founders inside AIPB
400kYouTube subscribers
38countries · live members
163kX followers
Members write up their wins in a 158-page doc — read it here →
Don't pick a favourite fish. Own the whole aquarium and route the job.
IX · the SOP

Wire Fugu into your stack this week.

1

Get a key at console.sakana.ai (or use the OpenRouter route if the console's not available to you).

2

Check your region first. One curl to the chat endpoint. A 403 on everything = the EU/UK wall, not a broken key.

3

Store the key in a file with tight permissions — never in a script, never in a repo.

4

Create the agent profile — one config pointing at the endpoint, key from the environment.

5

Run the verification test. A file-write job with a checkable end state. No agent joins the stack without passing it.

6

Give it a real job at default effort — xhigh. Only reach for max when a hard task underdelivers.

7

Watch the usage block. Orchestration tokens are the real bill. Note what a typical job costs before you scale it.

8

Route by lane. Fugu gets the big, hard, unattended builds. Your fast daily driver keeps the quick chats.

Thinking it?"What if Sakana changes access or pricing next month?"

Then you flip a different switch. That's the whole point of running models as profiles in one stack — Fugu is a staff member, not a foundation. Nothing in your workflow depends on any single vendor staying generous.

✦
X · the 30-day roadmap

From new API to staffed team member.

Week 1

Wire + verify

Key, profile, verification test. Fugu answers from inside your stack.

Week 2

Find its lanes

Run your real jobs through it at xhigh. Note where the expert team beats your daily driver.

Week 3

Unattended work

Give it the long builds overnight. Slow + good + unattended = free capacity.

Week 4

Cost tuning

Compare effort levels on your workload. Keep max for the jobs that earn it.

One prompt in. A team of experts out. That's the trade.
the recap

What you just gained.

i.
You understood the pufferfish.

One prompt, an orchestrator, up to three experts, one merged answer — and a separate bill for the team.

ii.
You saw it verified, not claimed.

A real agent run with a byte-exact receipt, wired into the Agent OS through one profile.

iii.
You got a playable data point.

8.6 on the bench's hardest open-world task, first attempt — embedded above, judged like every other model.

iv.
You know the gotchas.

The EU/UK region wall, ignored temperature, orchestration billing, and the 25-minute patience tax.

v.
You got the live scoreboard.

The rest of the 50-task run publishes itself — the link is the truth, not a screenshot.

vi.
You got the routing rule.

Big unattended builds go to the team. Quick chats stay fast. Nobody gets a monopoly.

Your move

Get the Sakana Fugu 1.1 setup — inside the whole operating system.

This guide gave you the wiring, the gotchas, and live playable proof. The Boardroom gives you the assembled machine: every frontier model as a switch, the bench harness, the workspace pattern, and the room that stress-tests new models the week they ship.

Readers watch launch threads. Operators had Fugu benched, wired and working the same day.

Decide which one you are tonight.

Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
4,000+ founders · 38 countries · live coaching calls every week · everything from this guide, pre-wired