The Goldie framework · K3 vs Fable 5 vs GPT-5.6 · live scores

The Frontier Shootout.

Three frontier models, identical one-shot build tasks, one tough judge — my GoldieBench leaderboard, not launch-thread vibes. GPT-5.6 Sol, Claude Fable 5, and this week's newcomer Kimi K3, with the same-task builds embedded so you can drive all three yourself.

Three colossal ornate machines facing each other across a moonlit arena — a brass mountain-machine, a golden owl-crowned engine and a silver orbital engine — while fully clothed winged referees in knee-length tunics hold scoring tablets
THE SHOOTOUT — three machines, one arena, playable receipts 3 contenders Sol · Fable 5 · K3 50 tasks identical one-shots One judge same rubric for all Podium 8.16 · 8.10 · 7.89every score clicks through to a playable build — check the judge yourself
the whole method in one line — same bullets, one referee, published evidence
50
identical tasks per model
1
judge, same rubric for all
8.16
the solo score to beat (GPT-5.6 Sol)
3.3×
price gap between the contenders
✦
I · the problem

The Launch-Thread Verdict Problem.

A new model drops and the verdicts arrive in minutes.

"Better than GPT." "Fable killer." "Benchmarks are rigged anyway."

None of those people ran the same task on all three models.

Vendor benchmarks compare different tests. Influencer takes compare screenshots. Reply guys compare feelings.

So you finish the thread knowing everyone's opinion and nobody's evidence.

Meanwhile the only comparison that matters for YOUR work — same job, same judge, same bar — never happens.

The Frontier Shootout breaks that cycle for good.

Thinking it?"Aren't all benchmarks gameable?"

Vendor-run ones, often. This one is mine: one-shot builds, real rendered screenshots, an independent judge model scoring the whole field on one rubric — and every build is published so you can check the judge.

When the evidence is playable, the argument gets short.

✦
II · how it works, in simple words

Same task. Same judge. Playable receipts.

GoldieBench is my model leaderboard. Every model gets the identical job list — one-shot, single-file builds: games, shaders, sims, tools. No retries, no hand-fixing.

Each build gets rendered for real, screenshotted, and scored 0–10 by the same judge model against the same rubric. Blank screens score 1. Task winners get medals.

This week Kimi K3 joined the board — same 50 tasks GPT-5.6 Sol and Claude Fable 5 already ran, and its full run is now complete: every launch-week rate-limit failure was re-run until all 50 tasks had a real judged build. What's left in its average are genuine results, including genuine wrecks.

·GPT-5.6 SolClaude Fable 5Kimi K3
Bench average8.16 / 10 (50 tasks)8.10 / 10 (47 tasks)7.89 / 10 (50 tasks — every one scored)
Task medals 🥇🥈🥉9 · 9 · 74 · 2 · 211 · 8 · 8 — most golds on the podium
Context window1.05M tokens200K (1M extended)1M tokens
Price per M (in/out)$5 / $30$10 / $50$3 in · or $0 on the Kimi coding plan
Personality on the benchconsistent all-roundershader/GPU-physics specialisthighest peaks, lowest floor — three 9.0s AND two broken builds

Live numbers from goldiebench.com — K3's full 50-task run is complete; click through for the current board.

GPT-5.6 Sol · 8.16 / 10 · 50 tasks scored Claude Fable 5 · 8.10 / 10 · 47 tasks scored Kimi K3 · 7.89 / 10 · all 50 tasks scored 0.27 separates first and third — but the medals tell a different story below
three frontier averages, to scale — 0.06 separates first and second
The live GoldieBench leaderboard: 851 live demos, 50 tasks, 22 models, 683 curated verdicts, with the ranked table below

What you're looking at: the actual board — 851 live demos, 50 tasks, 22 models. Every number in this guide clicks through to a playable build there. The two rows above 8.16 are multi-model ensembles (Fusion, Hermes MoA) — this shootout is the solo-model podium.

Thinking it?"Why is K3's average lower if it has more golds than Fable 5?"

Because averages punish failures and medals reward peaks. With all 50 tasks now scored, K3 holds ELEVEN golds — more than either rival — alongside two genuinely broken builds (Galaxy, Dogfight) that drag its average to third.

Highest peaks + lowest floor = read both numbers, not one.

TASK GOLDS — outright wins on identical briefs GPT-5.6 Sol 🥇 9 Kimi K3 · full run 🥇 11 Claude Fable 5 🥇 4 peaks ≠ averages: the week-old rookie now out-golds the entire podium
outright task wins — the THIRD-place model by average holds the MOST golds
✦
III · exactly how it works — 18 tests, three builds each

Same brief, three machines, eighteen times over.

Below: 18 of the 50 bench tasks, most impressive first — real games leading, then physics sims, shaders, and a working web desktop. Every cell is that model's ACTUAL one-shot build: the still is its real render, and one click loads the live, playable file. Scores and medals are the judge's, from the live board.

Nothing is hidden — including the wrecks. Where a score sinks to 3.x, that model shipped a genuinely broken build on that task. That's data too.

GAME

Voxelcraft

Minecraft-style sandbox — place and break blocks in a live world
Kimi K3 · 8.7🥉⛶ full screen
Claude Fable 5 · 9.0🥇⛶ full screen
GPT-5.6 Sol · 8.4⛶ full screen
GAME

Skyrim

First-person open-world fantasy explorer
Kimi K3 · 8.4⛶ full screen
Claude Fable 5 · 9.0🥇⛶ full screen
GPT-5.6 Sol · 8.4⛶ full screen
GAME

Dragonrealm

Frozen open world with walk-into encounters
Kimi K3 · 8.6🥉⛶ full screen
Claude Fable 5 · 7.8⛶ full screen
GPT-5.6 Sol · 8.6🥉⛶ full screen
GAME

Neonracer

Fullscreen neon racer with vapor-trail particles
Kimi K3 · 8.6🥇⛶ full screen
Claude Fable 5 · 8.3⛶ full screen
GPT-5.6 Sol · 8.6🥇⛶ full screen
GAME

Neoncity

Cyberpunk city you drive through
Kimi K3 · 8.7🥈⛶ full screen
Claude Fable 5 · 8.1⛶ full screen
GPT-5.6 Sol · 8.6🥉⛶ full screen
GAME

Outrun

Synthwave horizon driving game, pseudo-3D road
Kimi K3 · 8.6🥉⛶ full screen
Claude Fable 5 · 8.7🥇⛶ full screen
GPT-5.6 Sol · 8.7🥇⛶ full screen
GAME

Crypt

Torch-lit dungeon crawler
Kimi K3 · 7.4⛶ full screen
Claude Fable 5 · 8.8🥈⛶ full screen
GPT-5.6 Sol · 8.1⛶ full screen
GAME

Doom

Monsters in a raycaster maze that hunt you
Kimi K3 · 8.0⛶ full screen
Claude Fable 5 · 8.1⛶ full screen
GPT-5.6 Sol · 8.4⛶ full screen
GAME

Dogfight

Air-combat shooter
Kimi K3 · 3.5 (failed run)⛶ full screen
Claude Fable 5 · 7.4⛶ full screen
GPT-5.6 Sol · 8.6🥈⛶ full screen
SIM

Blackhole

Gravitational-lensing physics visualisation
Kimi K3 · 9.0🥇⛶ full screen
Claude Fable 5 · 8.7⛶ full screen
GPT-5.6 Sol · 8.8⛶ full screen
SIM

Fluid

WebGL fluid simulation, swirling particles
Kimi K3 · 9.0🥇⛶ full screen
Claude Fable 5 · 8.3⛶ full screen
GPT-5.6 Sol · 8.6🥉⛶ full screen
SIM

Pathtracer

Physically-correct ray-traced renderer
Kimi K3 · 8.6🥈⛶ full screen
Claude Fable 5 · 8.7🥇⛶ full screen
GPT-5.6 Sol · 6.4⛶ full screen
SIM

Galaxy

Particle galaxy you swirl with your mouse
Kimi K3 · 3.0 (failed run)⛶ full screen
Claude Fable 5 · 8.4⛶ full screen
GPT-5.6 Sol · 8.6🥇⛶ full screen
SIM

Boids

Emergent flocking-birds simulation
Kimi K3 · 8.6🥇⛶ full screen
Claude Fable 5 · 8.3⛶ full screen
GPT-5.6 Sol · 8.4⛶ full screen
VISUAL

Synthwave

Sunset-grid synthwave loop
Kimi K3 · 9.0🥇⛶ full screen
Claude Fable 5 · 8.6⛶ full screen
GPT-5.6 Sol · 8.7🥉⛶ full screen
VISUAL

Aurora

Northern-lights animation
Kimi K3 · 8.7🥇⛶ full screen
Claude Fable 5 · 8.6🥈⛶ full screen
GPT-5.6 Sol · 8.6🥈⛶ full screen
VISUAL

Waves

Animated ocean simulation
Kimi K3 · 8.7🥇⛶ full screen
Claude Fable 5 · 7.8⛶ full screen
GPT-5.6 Sol · 8.4⛶ full screen
PAGE

Webos

A working desktop OS — windows, dock, apps
Kimi K3 · 8.6🥉⛶ full screen
Claude Fable 5 · 8.0⛶ full screen
GPT-5.6 Sol · 8.4⛶ full screen

Every score above clicks through to a playable artifact — 54 builds on this page alone. This is what "comparison" should mean: not a table of adjectives.

Thinking it?"Why does K3 score 9.0 on one row and 3.0 a few rows later?"

Because both are real. K3's full run mixes outright task WINS (black hole, fluid, synthwave — three 9.0s, eleven golds total) with genuine faceplants (Galaxy, Dogfight — real builds that shipped broken). Highest peaks, lowest floor — exactly why you read medals AND averages, never one number.

Boards that hide the wrecks aren't benchmarks; they're brochures.

Want the other 32 tasks? Every head-to-head is on the bench — K3 vs GPT-5.6, K3 vs Fable 5, Fable 5 vs GPT-5.6 — every build playable.

IV · the sources

The receipts, all clickable.

"🚨 BREAKING: Kimi K3 benchmarks have released and it's ~Fable/Sol level"

— leo (@synthwavedd), on X, 16 July 2026 — the claim this guide puts to the test

The claim · now testable

"~Fable/Sol level" — is it?

This launch-day post set the frame: K3 belongs in the Fable/Sol tier. On my board the claim holds up with an asterisk — K3 took MORE task golds than Sol or Fable 5 (eleven), but two broken builds pin its average to third at 7.89. Peak-tier: yes. Reliability-tier: not yet.

Kimi K3's GoldieBench scorecard page: 1,048,576-token context, 3 dollars per million input, 50 tasks tested, 9 golds 5 silvers 7 bronzes, released 2026-07-16

What you're looking at: K3's live scorecard on the board — the 1M context, the $3/M pricing, 50 tasks tested and the day-one medal haul, one day after release.

Everything behind this shootout ↓
When the evidence is playable, the argument gets short.
V · the framework

The five rounds of the Frontier Shootout.

THE INVOICE COLUMN — input price per million tokens Claude Fable 5 · $10 / M in · $50 / M out GPT-5.6 Sol · $5 / M in · $30 / M out Kimi K3 · $3 / M in — or $0 on the Kimi coding plan → $0 on plan
same scale, real list prices — a 3.3× spread between the top two contenders and the rookie
i.

Same Bullets

Identical prompts for every contender — one-shot, no retries, no favours. Different tests can't be compared; identical ones can't be argued with.

ii.

One Referee

A single judge model scores the whole field on one rubric, looking at real rendered screenshots. Consistency beats sophistication.

iii.

Published Evidence

Every build ships to the site, playable. A score you can't check is a rumour with decimals.

iv.

Peaks AND Floors

Read medals for capability, averages for reliability. K3's nine golds and its climbing average are both true — lanes come from both.

v.

Price In The Frame

8.16 at $5/$30, 8.10 at $10/$50, high-peaks at $3-or-$0. The score without the invoice is half a comparison.

THE FRONTIER SHOOTOUT — five rounds, one verdict i · SAMEBULLETS ii · ONEREFEREE iii · PUBLISHEDEVIDENCE iv · PEAKSAND FLOORS v · PRICE INTHE FRAME out the other end: a routing table, not a favourite — and it reruns itself every launch
the whole framework in one line — every round removes a way to fool yourself
✦
VI · old way vs new way

Picking a model, two ways.

Old way
hours of threads, zero evidence
  • Read forty launch-thread verdicts from people who ran nothing
  • Compare vendor benchmarks that test different things
  • Switch your whole stack to whatever won the news cycle
  • Discover its weaknesses on YOUR work, in production
  • Repeat next launch, from zero
  • Never once see the same task run on all contenders
New way
one board, playable proof, per-job routing
  • Identical tasks, one judge, whole field — automatically per launch
  • Open the same-task builds side by side and drive them
  • Read peaks (medals) and floors (averages) separately
  • Divide by price: Sol for balance, Fable for shaders, K3 for long context at $0-on-plan
  • Route jobs by lane; keep every model that wins one
  • Next launch: same harness, evidence by dinner
Thinking it?"Just tell me which one to use."

By the current board: GPT-5.6 Sol if you want one all-rounder; Fable 5 if your work leans visual/shader-heavy and budget allows; K3 if million-token context or plan-included pricing is the constraint — with its ranking still settling.

But the honest answer is the framework: route by lane, not by loyalty.

Skip the setup

Get the Frontier Shootout harness built for you.

You can eyeball the public board free, forever. The Boardroom gives you the machinery behind it — the Agent OS where all three contenders are wired as switches, plus the bench harness to score them on YOUR tasks.

The full Agent OS zip — GPT-5.6, Fable 5 and K3 all pre-wired as profiles
The bench harness — run your own shootout on your own briefs
Coaching calls where we route your real workload by lane
A room of 3,900+ operators posting real model results daily
The prompts, the SOPs, and a member map for your city
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Set up in an afternoon · used in 38 countries · new models added the week they ship
Thinking it?"Doesn't running the Agent OS burn a fortune in tokens?"

No — that's the biggest myth about it. The everyday 90% runs on free local models on your own machine, and free APIs slot in for more — one of this guide's three contenders is included in a coding plan many members already pay for.

For the frontier work, the Agent OS drives the plans and CLIs you already own — Claude's CLI with your Claude subscription, K3 with the Kimi plan. A layer on top, never a second meter.

And inside the AI Profit Boardroom there are full token-optimisation tutorials, so usage drops even further.

✦
VII · three beliefs to drop

What's actually holding you back.

Wrong: "There's one best model and the game is finding it."

Right: The board shows three different winners depending on the lane — balance, shaders, long context. The game is routing, and routers beat loyalists every month.

Wrong: "Day-one scores are final scores."

Right: K3's published average climbed from 5.81 on launch day to 7.89 once every failure was properly re-run. Launch-day numbers are opening bids — trust boards that keep re-running, not screenshots of hour one.

Wrong: "Benchmarks don't transfer to my real work anyway."

Right: Generic ones don't have to — the same harness runs on YOUR briefs. Fifty of my real tasks judged consistently beats any public leaderboard for my routing decisions, and the harness is reusable.

Don't take my word for it

Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.

Read the 158-page wins doc →
✦
VIII · the receipts

Who's already running this.

GoldieBench is public — 29 models, 50 tasks, over a thousand published builds, every score checkable against a playable artifact. The three-way table above links straight into it.

3,900+founders inside AIPB
400kYouTube subscribers
38countries · live members
163kX followers
Members write up their wins in a 158-page doc — read it here →
✦
IX · the SOP

Run your own shootout this week.

1

Pick five of YOUR real briefs. The page you'd build, the script you'd write, the audit you'd run — not abstract puzzles.

2

Freeze the prompts. Word-identical for every model. The moment you tailor per model, you're benchmarking your prompting.

3

One shot each, no rescues. Save every output as-is. Failures are data, not embarrassments.

4

Judge blind and consistently. Same rubric, ideally a judge model, scores 0–10. Never grade your favourite differently.

5

Record peaks and floors separately. Best-task wins tell you capability; averages tell you trust.

6

Add the invoice column. Score per dollar changes podiums — a 7.9 at $0 beats an 8.1 at $50 for plenty of lanes.

7

Route, don't crown. Each model keeps the lanes it won. Loyalty is for sports teams.

8

Re-run on every launch. The harness is the asset — next model, same five briefs, evidence by dinner.

✦
X · the 30-day roadmap

From thread-reader to referee.

Week 1

Wire all three

Sol, Fable and K3 as switches in one stack — cheapest routes first (plans you own).

Week 2

Your five briefs

Freeze them, run the shootout, judge blind. First real routing table.

Week 3

Route production

Send each lane's work to its winner for a week. Measure what changed.

Week 4

Automate the referee

Script the harness so the next launch costs you an evening, not a debate.

Route by lane, not by loyalty. The board decides.
the recap

What you just gained.

i.
You gained real numbers.

8.16 vs 8.10 vs 7.89 — same 50 tasks, same judge, live board.

ii.
You gained playable proof.

The same brief built by all three, embedded — drive the evidence yourself.

iii.
You gained the two-number habit.

Peaks for capability, floors for trust. One number lies; two triangulate.

iv.
You gained the price column.

$5/$30 vs $10/$50 vs $3-or-$0. Podiums move when the invoice shows up.

v.
You gained lane routing.

Sol all-round, Fable for shader-visual work, K3 for million-token jobs.

vi.
You gained a reusable referee.

The 8-step shootout runs on your briefs, every launch, forever.

Your move

Run the Frontier Shootout — or keep outsourcing your verdicts to reply guys.

This guide gave you the live scores, the playable three-way, and the harness recipe. The Boardroom gives you the machine — all three contenders wired into one Agent OS, the bench harness, and a room of operators comparing real results daily.

Readers pick a favourite from a thread. Operators run five briefs and know by Friday.

Decide which one you are tonight.

Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
3,900+ founders · 38 countries · live coaching calls every week · everything from this guide, pre-wired