One prompt. Three sentences long.
And Claude built a fully playable 3D game at the level of a big-budget console title.
Fifty-five thousand lines of code — every sound, every texture, every 3D model made by the AI itself.
Nothing downloaded. Nothing borrowed.
The AI ran on its own until the game existed.
He calls it the Gauntlet Loop — and in two weeks people have used it for racing games, 3D property walkthroughs, landing pages and social media designs.
Matt didn't just show the game — he shipped a free tool that writes the Gauntlet Loop prompt for whatever you want to build. This is the gauntlet-built game from his post, playing on loop.
The gauntlet-built game demo — video by @mattshumer_ · watch the original post on X ↗
Builders ran with it immediately. Here's one taking a single panorama photo and one prompt — and getting an entire explorable 3D town out, exportable to real game engines.
A whole 3D town from one panorama and one prompt — video by @HomeFreeHatGuy · watch the original post on X ↗
What you're watching: Matt's published prompt, retyped into Claude Code — the task, the method, the bar. Three sentences in; a fully playable 55,000-line game out.
Here's how everyone uses AI right now.
You send a prompt. The AI gives you something.
You look at it. You say "no, fix this part."
Still not right. Twenty, thirty, fifty rounds of this.
You are the quality checker — every draft has to pass through your eyeballs.
Which means the AI can only work as fast as you can review.
It was — until the checking got handed to AI too. That handover is this whole guide.
"I don't prompt Claude anymore. My job is to write loops."
— Boris Cherny, creator of Claude Code at Anthropic, June 2026
Anthropic wrote this up back in 2024: work comes out better when a different model plays the evaluator.
"One LLM call generates a response while another provides evaluation and feedback in a loop."
— Anthropic, "Building Effective Agents", December 2024
Running the gauntlet means you face every judge in the line — and you have to survive every single one to get through. Three upgrades make it work.
A lead agent breaks your goal into small pieces and fans out subagents — one owned lighting, one the vehicles, one the sound, one the physics. AI does small, focused jobs far better than giant vague ones.
The critics never see the code or the builder's excuses — only the rendered result, compared against something real. A critic that watches the builder work starts sympathising with it; a blind critic sees the result cold, the way a customer would.
What you're watching: a real builder-vs-critic round we ran in Claude Code for this guide — the builder ships a landing page, a blind critic judges only the screenshot, FAILS it 6.0/10 with five specific fixes, and round two comes back stronger. Real session, replayed at reading pace.
Matt's critics compared the AI's game against actual screenshots of the real title, and the stop condition was brutal: don't stop until each critic is utterly wowed.
Because here's the dirty secret: left alone, AI always calls its own work done.
A July research paper watched an agent run 54 loop cycles — it claimed improvement in every single one, but measured results got worse or stayed flat over half the time.
That's the disease. The gauntlet loop is the cure.
A wall of judges that never gets tired, never gets bored, never gets polite, and never says "eh, good enough." Human reviewers burn out by draft three — these critics reject draft two hundred with the same cold energy as draft one.
"No one in their right mind would ever spend the time to write something this custom. But AI models have all the stamina and patience in the world."
— Andrej Karpathy, former Tesla AI director, on hyper-custom AI-built worlds, Aug 2026
The full picture: goal in → lead agent splits it → subagents build → blind critics judge every piece against a real benchmark → failures cycle back — until the work survives the entire gauntlet.
Want this actually running on your business — not just understood? The Agent OS inside the AI Profit Boardroom has the gauntlet pattern wired in. Plug in your Claude, your Hermes, your OpenClaw, your Free Claude Code — and run builder-vs-critic gauntlets on real business assets.
That's the biggest myth about it. The everyday 90% runs on free local models on your own machine, free APIs slot in for more, and for frontier work it drives the CLIs you already pay for — your Claude subscription already includes the Claude Code CLI, so you're never paying twice.
And inside the Boardroom there are full token-optimisation tutorials, so usage drops to the bone and stays there.
One developer changed one word of Matt's prompt — Formula-1-style racing instead of the original game. Before building a single road, the gauntlet built a screenshot tool, a lighting reader, a round-comparison tool, and 136 more — including one that drove the car through a scripted route so the critics could watch it race.
The gauntlet knew its own ceiling. That honesty came from the blind critics — a builder alone would have declared victory on day one.
Here's a builder running his own version of the loop on an existing game just to raise its quality — and shipping the playable result. The loop doesn't care if the work is new or old; it just keeps rejecting until the critics are happy.
An old game, quality-raised by a gauntlet run — video by @LexnLin · watch the original post on X ↗
Another run rebuilt a classic open-world racer — and invented the best picture of how gauntlet critics think: it cut real screenshots into grids and rebuilt the world square by square until each square passed its judge.
One test fed a gauntlet a real apartment floor plan plus room photos: two hours later, a walkable 3D version — marble kitchen counter, the painting on the wall — with the loop generating its own real-photo-vs-my-version progress report and failing its own rounds.
A design gauntlet pointed at a proven 15,000-like carousel spun up three blind critics — the brief, the design system, the craft — and caught missing logos, an undersized brand mark, and a stray lime accent across ten rounds.
Carousels, presentations, landing pages, proposals, sales pages — a build team and a panel of judges arguing in a loop while you do something else.
Wrong: "This is for developers. I could never run a gauntlet loop."
Right: It's a prompt structure, not a coding skill — three sentences: the task, the method, the bar. If you can describe the outcome and screenshot a benchmark, you can run one. The skill that matters now isn't writing code — it's writing standards.
Proof of how light this really is: one builder packaged the entire Gauntlet Loop method as a single markdown file you install into your coding agent with one command. If it fits in one text file, it's not a developer skill — it's a writing skill. Here's the kind of game that one file produces.
Built with the one-file gauntlet skill — demo shared in @arjunkshah21's post (gameplay by @0xRishi) · watch on X ↗
Wrong: "A loop that never stops must eat my whole plan."
Right: Design gauntlets on business assets run around 2–3 million tokens — and the lever most people miss is that the subagents don't need to be premium. One documented run used a budget model for the whole gauntlet and built a complete course website plus a full testing suite in 1 hour 13 minutes, on about 1% of a daily usage limit. Claude as the judge, cheap models as the workers — inside the Agent OS you swap models per agent.
Wrong: "AI work is never actually good enough, so why bother."
Right: That belief used to be correct — because the AI checked its own work and called it finished. The gauntlet removes the AI's right to grade itself: your taste becomes a checklist, the blind critics enforce it round after round, and the first draft you ever see has already survived ten rounds of rejection.
Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.
Read the 158-page wins doc →"Utterly wowed" has no finish line — that's by design. Matt's own published scores: the blind critics started his game at 3.6/10, days of looping pushed it to 5.1, and in every blind round every critic still picked the real game. He pulled the plug himself. Set a round limit — you are the final judge, and that job stays human.
"Make it amazing" is not a bar. Take the reference away and the blind critics have nothing to measure with — slowly, quietly, they start agreeing with the builder. One documented landing-page run with no strong reference produced a page that looked genuinely beautiful and matched the actual brand's style not at all. Ten rounds of polishing, pointed in the wrong direction. The gauntlet is a polisher, not a compass — you point it.
Screenshots match or they don't. Tests pass or they fail. A number moves or it doesn't. When the finish line is checkable this handles serious work: PostHog pointed an overnight loop at their query engine and found a bug that had sat there for three years, and Anthropic ran 16 agents across ~2,000 sessions and built a working C compiler that compiles the Linux kernel. Checkable finish line: the loop earns every token. Vibes finish line: the loop wanders.
What you're watching: the three sentences being written into Claude Code for a real business asset — a carousel — and the loop accepting the job. This is the whole ten minutes.
Running it with no real benchmark. That's the mistake most people copying this prompt are making right now — the critics end up with nothing to measure against, agree with the builder, and polish in the wrong direction for hours. Screenshot the best version you've ever seen FIRST. That's the whole game.
For your entire working life, checking was the expensive part.
Someone had to look at every draft — and that someone was you.
Now a wall of blind critics can inspect your work 1,329 times overnight without a single sigh.
The business owners who learn to write gauntlets first will ship better work with less of their own time in the middle.
The ones still prompting one message at a time will quietly wonder why everything they make feels slower and rougher than the competition's.
Readers bookmark this page and keep prompting one message at a time. Operators join, wire in the gauntlet pattern this week, and let a wall of critics polish the pages, content and offers that actually bring them customers.
The people winning with gauntlet loops aren't smarter than you.
They stopped judging every draft themselves and wrote the standard once instead.
Three sentences. One benchmark image. Walk away.
The prompt is public. The pattern is open.
The only question is whether you set your bar this week — or watch someone in your market set theirs first.