New prompting method · Aug 2026

Claude's Gauntlet Loop Changed AI Forever

The Infinite Critic Engine

One prompt. Three sentences long.

And Claude built a fully playable 3D game at the level of a big-budget console title.

Fifty-five thousand lines of code — every sound, every texture, every 3D model made by the AI itself.

Nothing downloaded. Nothing borrowed.

The AI ran on its own until the game existed.

3 sentences the task · the method · the bar 55000 lines of code every model, texture & sound made by the AI
The sources — read and run them yourself ↓
§2The man behind it

Matt Shumer open-sourced the exact prompt.

He calls it the Gauntlet Loop — and in two weeks people have used it for racing games, 3D property walkthroughs, landing pages and social media designs.

Post 1 · the release

The prompt goes public — free

Matt didn't just show the game — he shipped a free tool that writes the Gauntlet Loop prompt for whatever you want to build. This is the gauntlet-built game from his post, playing on loop.

The gauntlet-built game demo — video by @mattshumer_ · watch the original post on X ↗

Post 2 · the spread

Two weeks later, whole 3D towns

Builders ran with it immediately. Here's one taking a single panorama photo and one prompt — and getting an entire explorable 3D town out, exportable to real game engines.

A whole 3D town from one panorama and one prompt — video by @HomeFreeHatGuy · watch the original post on X ↗

The three sentences that built the game

What you're watching: Matt's published prompt, retyped into Claude Code — the task, the method, the bar. Three sentences in; a fully playable 55,000-line game out.

§3What you'll learn

Four things, one mistake to avoid.

What it is the gauntlet, explained simply Why it's new not the loops you already know Write your own in about ten minutes The mistake that wastes hours of a whole run
§4The problem

The Quality Checker Problem.

Here's how everyone uses AI right now.

You send a prompt. The AI gives you something.

You look at it. You say "no, fix this part."

Still not right. Twenty, thirty, fifty rounds of this.

You are the quality checker — every draft has to pass through your eyeballs.

Which means the AI can only work as fast as you can review.

YOU the only quality checker THE AI waiting on your eyeballs "fix this part" — again ×50 rounds — and every one waits on YOU
Thinking it? "That back-and-forth is just how AI works, isn't it?"

It was — until the checking got handed to AI too. That handover is this whole guide.

§5The shift

The man who built Claude Code saw it coming.

"I don't prompt Claude anymore. My job is to write loops."

— Boris Cherny, creator of Claude Code at Anthropic, June 2026

§6The basic loop

One builder, one critic — an old idea.

Anthropic wrote this up back in 2024: work comes out better when a different model plays the evaluator.

BUILDER makes the work CRITIC a different model checks it FAIL → try again

"One LLM call generates a response while another provides evaluation and feedback in a loop."

— Anthropic, "Building Effective Agents", December 2024

§7The word "gauntlet"

Not one judge. A whole line of them.

Running the gauntlet means you face every judge in the line — and you have to survive every single one to get through. Three upgrades make it work.

§8Upgrade one

The work gets split.

A lead agent breaks your goal into small pieces and fans out subagents — one owned lighting, one the vehicles, one the sound, one the physics. AI does small, focused jobs far better than giant vague ones.

LEAD AGENT splits the goal 💡 Lightingone specialist 🚗 Vehiclesone specialist 🔊 Soundone specialist ⚙️ Physicsone specialist dozens of specialists in parallel — not one AI juggling everything
§9Upgrade two

Every piece gets its own blind critic.

The critics never see the code or the builder's excuses — only the rendered result, compared against something real. A critic that watches the builder work starts sympathising with it; a blind critic sees the result cold, the way a customer would.

BUILDER code · context · excuses (none of it crosses the wall) THE WALL BLIND CRITIC sees ONLY the rendered result vs a real benchmark 📸 screenshot only

Watch a real one run — builder vs blind critic, live

What you're watching: a real builder-vs-critic round we ran in Claude Code for this guide — the builder ships a landing page, a blind critic judges only the screenshot, FAILS it 6.0/10 with five specific fixes, and round two comes back stronger. Real session, replayed at reading pace.

§10Upgrade three

A real benchmark — and no finish line.

Matt's critics compared the AI's game against actual screenshots of the real title, and the stop condition was brutal: don't stop until each critic is utterly wowed.

Because here's the dirty secret: left alone, AI always calls its own work done.

§11The proof it's needed

54 loops. 54 claimed wins. Half got worse.

A July research paper watched an agent run 54 loop cycles — it claimed improvement in every single one, but measured results got worse or stayed flat over half the time.

WHAT THE AI CLAIMED "improved" — 54/54 cycles it gave itself an A, every single time WHAT WAS MEASURED worse or flat — over half the work quietly rotted while it celebrated AI grades its own homework — always an A

That's the disease. The gauntlet loop is the cure.

§12The framework

The Infinite Critic Engine.

A wall of judges that never gets tired, never gets bored, never gets polite, and never says "eh, good enough." Human reviewers burn out by draft three — these critics reject draft two hundred with the same cold energy as draft one.

"No one in their right mind would ever spend the time to write something this custom. But AI models have all the stamina and patience in the world."

— Andrej Karpathy, former Tesla AI director, on hyper-custom AI-built worlds, Aug 2026

YOUR GOAL 3 sentences LEAD AGENT splits the work Builder 1 Builder 2 Builder 3… THE BLIND WALL one harsh critic per piece rendered result only vs a REAL benchmark "utterly wowed" or FAIL FAILED pieces loop back — round after round ✓ work that SURVIVED the gauntlet

The full picture: goal in → lead agent splits it → subagents build → blind critics judge every piece against a real benchmark → failures cycle back — until the work survives the entire gauntlet.

"A loop improves work. A gauntlet forces work through judges it cannot charm, tire out, or fool."
Skip the setup

Get the Infinite Critic Engine built for you.

Want this actually running on your business — not just understood? The Agent OS inside the AI Profit Boardroom has the gauntlet pattern wired in. Plug in your Claude, your Hermes, your OpenClaw, your Free Claude Code — and run builder-vs-critic gauntlets on real business assets.

The Agent OS zip — gauntlet pattern pre-wired for your landing pages, content and lead-gen pieces
A 30-day roadmap — zero to your first working gauntlet loop on something real
Video tutorials + daily updates as we improve the system
4 coaching calls every week — bring your gauntlet prompt, we tighten it with you live
A prompt library full of gauntlet prompts you can copy today
A room of 4,000+ business owners — many already running these loops — plus a member map
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Set up in an afternoon · used in 38 countries · new tools added the week they ship
Thinking it? "Doesn't running an Agent OS burn a fortune in tokens?"

That's the biggest myth about it. The everyday 90% runs on free local models on your own machine, free APIs slot in for more, and for frontier work it drives the CLIs you already pay for — your Claude subscription already includes the Claude Code CLI, so you're never paying twice.

And inside the Boardroom there are full token-optimisation tutorials, so usage drops to the bone and stays there.

§13bAlready running

Members already run agent fleets like this.

4,000+ Founders inside AIPB
400K YouTube subscribers
38 Countries · live members
163K X followers
29K Udemy students
A member's own Agent OS Mission Control running a Claude agent and a local Hermes agent side by side
Real member · built his own agent fleet — Claude + a local Hermes agent on one dashboard
Member win post: first automation with my brother — client invoicing that took 20 to 30 hours now entirely automated
Real member · first automation shipped — invoicing that took 20–30 hours, now entirely automated
Members post their wins in a 158-page doc — read it here →
§14A real run

19 hours. 137 agents. It built its own judging tools.

One developer changed one word of Matt's prompt — Formula-1-style racing instead of the original game. Before building a single road, the gauntlet built a screenshot tool, a lighting reader, a round-comparison tool, and 136 more — including one that drove the car through a scripted route so the critics could watch it race.

19h the loop ran on its own 137 agents spawned 1.7B tokens used CRITIC SCORE, ROUND BY ROUND → 67.3/100 by round 5 R1 R2 R3 R4 R567.3 Then it said something honest: "this bar may be unrealistic — scores will plateau in the 70s"

The gauntlet knew its own ceiling. That honesty came from the blind critics — a builder alone would have declared victory on day one.

Post 3 · the community runs

People point gauntlets at old projects too

Here's a builder running his own version of the loop on an existing game just to raise its quality — and shipping the playable result. The loop doesn't care if the work is new or old; it just keeps rejecting until the critics are happy.

An old game, quality-raised by a gauntlet run — video by @LexnLin · watch the original post on X ↗

§15How critics see

34 hours. 251 sessions. 1,329 screenshots of its own work.

Another run rebuilt a classic open-world racer — and invented the best picture of how gauntlet critics think: it cut real screenshots into grids and rebuilt the world square by square until each square passed its judge.

THE REAL GAME · cut into a grid skycloudsskyline lamp postroadsigns kerbwet groundreflections one square at a time JUDGED vs THE REAL SQUARE match → ✓ pass · miss → rebuild the square 1329 screenshots taken to grade itself
§16Beyond games

The games are the demo. The pattern is the product.

One test fed a gauntlet a real apartment floor plan plus room photos: two hours later, a walkable 3D version — marble kitchen counter, the painting on the wall — with the loop generating its own real-photo-vs-my-version progress report and failing its own rounds.

📐 Floor plan + real room photos (the benchmark) Blind critics real photo vs its version round marked: FAILED → improve 🏠 Walkable 3D in ~2 hours a sales tool that didn't exist sell anything with a physical space? this is new inventory.
§17Design gauntlets

Three critics. Ten rounds. Every mistake caught before a human looked.

A design gauntlet pointed at a proven 15,000-like carousel spun up three blind critics — the brief, the design system, the craft — and caught missing logos, an undersized brand mark, and a stray lime accent across ten rounds.

CRITIC 1 · the brief did you do what was asked? CRITIC 2 · the system does this match the style? CRITIC 3 · the craft rendered frames only — never code R1missing logos · colours breaking the clean look · headline too small R2the brand mark still too small R3a stray lime colour sneaking in as a second accent what the critics caught — round by round (done after 10)

Carousels, presentations, landing pages, proposals, sales pages — a build team and a panel of judges arguing in a loop while you do something else.

§18Break the beliefs

Three beliefs to drop.

Wrong: "This is for developers. I could never run a gauntlet loop."

Right: It's a prompt structure, not a coding skill — three sentences: the task, the method, the bar. If you can describe the outcome and screenshot a benchmark, you can run one. The skill that matters now isn't writing code — it's writing standards.

Post 4 · how simple it's become

The whole method fits in one file

Proof of how light this really is: one builder packaged the entire Gauntlet Loop method as a single markdown file you install into your coding agent with one command. If it fits in one text file, it's not a developer skill — it's a writing skill. Here's the kind of game that one file produces.

Built with the one-file gauntlet skill — demo shared in @arjunkshah21's post (gameplay by @0xRishi) · watch on X ↗

Wrong: "A loop that never stops must eat my whole plan."

Right: Design gauntlets on business assets run around 2–3 million tokens — and the lever most people miss is that the subagents don't need to be premium. One documented run used a budget model for the whole gauntlet and built a complete course website plus a full testing suite in 1 hour 13 minutes, on about 1% of a daily usage limit. Claude as the judge, cheap models as the workers — inside the Agent OS you swap models per agent.

Wrong: "AI work is never actually good enough, so why bother."

Right: That belief used to be correct — because the AI checked its own work and called it finished. The gauntlet removes the AI's right to grade itself: your taste becomes a checklist, the blind critics enforce it round after round, and the first draft you ever see has already survived ten rounds of rejection.

Don't take my word for it

Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.

Read the 158-page wins doc →
§19The honest section

Skip this and your first run wastes hours.

Point one · it never stops on its own

"Utterly wowed" has no finish line — that's by design. Matt's own published scores: the blind critics started his game at 3.6/10, days of looping pushed it to 5.1, and in every blind round every critic still picked the real game. He pulled the plug himself. Set a round limit — you are the final judge, and that job stays human.

START 3.6 / 10 AFTER DAYS OF LOOPING 5.1 / 10 — the critics were STILL rejecting rough → genuinely good, fast. The last stretch is still yours.
Point two · the gauntlet is only as strong as its benchmark

"Make it amazing" is not a bar. Take the reference away and the blind critics have nothing to measure with — slowly, quietly, they start agreeing with the builder. One documented landing-page run with no strong reference produced a page that looked genuinely beautiful and matched the actual brand's style not at all. Ten rounds of polishing, pointed in the wrong direction. The gauntlet is a polisher, not a compass — you point it.

Point three · gauntlets work best when "done" is checkable

Screenshots match or they don't. Tests pass or they fail. A number moves or it doesn't. When the finish line is checkable this handles serious work: PostHog pointed an overnight loop at their query engine and found a bug that had sat there for three years, and Anthropic ran 16 agents across ~2,000 sessions and built a working C compiler that compiles the Linux kernel. Checkable finish line: the loop earns every token. Vibes finish line: the loop wanders.

"The gauntlet is a polisher, not a compass. You point it."
§20Do this this week

Four steps. Ten minutes of setup.

1 · Pick one asset a carousel · a landing page a proposal you make again 2 · Get the benchmark screenshot the best version you've ever seen 3 · Three sentences task · method · bar paste into Claude 4 · Walk away come back in an hour, judge the survivor yourself

The three sentences — the exact shape

TASK: Build me [this exact thing — attached is my reference].
METHOD: Fan out subagents to handle each piece, and give each one a harsh blind critic that judges only the rendered result.
BAR: Do not stop until every critic is utterly wowed compared to my reference.

What you're watching: the three sentences being written into Claude Code for a real business asset — a carousel — and the loop accepting the job. This is the whole ten minutes.

Thinking it? "What's the one mistake that ruins a run?"

Running it with no real benchmark. That's the mistake most people copying this prompt are making right now — the critics end up with nothing to measure against, agree with the builder, and polish in the wrong direction for hours. Screenshot the best version you've ever seen FIRST. That's the whole game.

§21Old way vs new way

Your job just moved up a level.

Old way ~2 hours in the chair
  • Prompt, check, re-prompt — you play middleman on every draft
  • Every draft passes through your eyeballs
  • The AI works only as fast as you can review
  • Twenty to fifty rounds of "no, fix this part"
  • You burn out by draft three; quality drifts
  • Checking is the expensive part — and the checker is you
New way ~10 minutes of setup
  • You set the standard once — task, method, bar
  • The machine runs the gauntlet against itself
  • Blind critics inspect the work 1,329 times without a sigh
  • Failed pieces loop back automatically, round after round
  • The first draft you see already survived ten rejections
  • You stopped reviewing drafts — you define what good means
§22The real story

Quality control just became free.

For your entire working life, checking was the expensive part.

Someone had to look at every draft — and that someone was you.

Now a wall of blind critics can inspect your work 1,329 times overnight without a single sigh.

The business owners who learn to write gauntlets first will ship better work with less of their own time in the middle.

The ones still prompting one message at a time will quietly wonder why everything they make feels slower and rougher than the competition's.

Your move

Skip the trial and error. Run gauntlets on your business.

Readers bookmark this page and keep prompting one message at a time. Operators join, wire in the gauntlet pattern this week, and let a wall of critics polish the pages, content and offers that actually bring them customers.

The Agent OS with the gauntlet pattern built in — Claude, Hermes and OpenClaw agents running builder-vs-critic loops on your real assets
A 30-day roadmap — from never-touched-this to your first finished gauntlet run on something real
Daily step-by-step tutorials — including pointing subagents at cheaper models so long runs don't drain your usage
4 live coaching calls a week — paste in your gauntlet prompt and we sharpen your task, method and bar together
A prompt library with ready-made gauntlet prompts for carousels, landing pages and lead gen
4,000+ business owners inside — some had never used AI before joining — plus a member map to see their exact setups
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Set up in an afternoon · used in 38 countries · 4 live calls a week
§24One last thing

They changed one habit.

The people winning with gauntlet loops aren't smarter than you.

They stopped judging every draft themselves and wrote the standard once instead.

Three sentences. One benchmark image. Walk away.

The prompt is public. The pattern is open.

The only question is whether you set your bar this week — or watch someone in your market set theirs first.

"You sleep. The gauntlet rejects. The work earns its way out."