Hands-on test · 30 Jun 2026 · run on my own M4 Max

The Local AI Engine™. I ran the #1 trending local model. It blanked.

I downloaded the most-hyped new local model on Hugging Face — a 12B coder "distilled from Fable 5" — and pointed it at my flagship build test. It's fast. It also handed me a black screen — at first. Then the engine iterated with it until it shipped a real, walkable world. Here's the honest before and after.

A robed scholar fully clothed in a long flowing hooded chiton covering the whole body, standing at a dark stone proving-ground table in a frozen mountain dragon realm at dusk, lifting one glowing crystalline orb toward the light while a row of other orbs sit dark and cold beside it — testing which models still glow
GoldieBench · "The Dragon Realm" flagship build · run locally on an M4 Max
gemma-4-12B · raw one-shotthe trendy one · ~37 tok/s
black screen
gemma-4-12B · + the enginesame model, iterated to working
walkable ✓
Qwable 27B · the proven onealready in my system · ~18 tok/s
first try ✓
Same prompt, same machine. Raw, the trendy 12B rendered a black screen. The engine iterated with it until it shipped a walkable world — and the proven 27B nailed it first try. The model was never the moat; the system is.
0
tok/s on my Mac (fast)
0
prompt → playable game (after the engine)
0
GB on disk — runs light
$0
cost — 100% local
Run it yourself — the real sources ↓
✦
My story · why this test matters

I used to chase every shiny local model. Then I built the engine.

Before

Every week a new "best" local model trends on Hugging Face.

I'd see the hype, download it, and feel behind until I had it running.

I'd point it at a real job — build me a page, build me a game — and half the time it gave me a blank screen or broken code.

I'd lose a whole evening before I even knew if the model was any good.

And the second I got it set up, a newer one dropped and I started over.

Then I stopped trusting models and started testing them.

After

Now every new model drops into one engine on my Mac.

I run it on the same real build test in minutes — and the system tells me, with a screenshot, whether it actually works.

When the trendy one blanks, a proven model takes the job instead, automatically.

I stopped losing evenings to hype. I keep the models that ship and drop the ones that don't.

You can run this same engine. Same Mac. Same test.

the receipts

Real operators. Real builds. Inside the Boardroom right now.

I'm not the only one running models locally and shipping with them. Here's the room you'd be joining — agency owners, ecom founders, course creators, solo builders, across 38 countries.

3,900+Founders inside AIPB
400kYouTube subscribers
38Countries · live members
163kX / Twitter followers

I'm not going to paste invented quotes here. The wins are real and written by the members themselves — agency owners, ecom founders, course creators, solo operators across 38 countries. Read them in their own words.

Read the 158-page wins doc →
Before you scroll on —

Commit to testing, not trusting.

You've seen the matchup at the top. Fast model. Black screen.

The next ten minutes show exactly how I test a model in minutes instead of losing an evening.

So here's the deal.

Promise yourself one thing right now. Before you download the next trendy model, you'll put it through a real build test first — and judge it on what renders, not what trends. Because the moment you make that switch, you stop wasting nights on hype.

The people chasing the model of the week are running in circles. The people who built a system test once and move on.

Be one of those people.

Commit to testing every model today. It changes how you pick tools forever.

✦
I · the framework

The Local AI Engine™.

Here's the simple idea behind everything below.

You don't pick the perfect model. You build an engine that runs any model, tests it honestly, and keeps the ones that ship. The model is just fuel. The engine is what you own.

i.

The Swap

Any local model drops into one slot on your Mac. New one trends? Swap it in, no rewiring. You never depend on a single model again.

ii.

The Test

Every model runs the same real build, one shot. You judge it on what renders — a screenshot — not on a benchmark number or a hype thread.

iii.

The Catch

The engine spots the blank screens and broken builds before you ever ship them. The bad output dies quietly instead of going to a client.

iv.

The Fallback

When the trendy model blanks, a proven model takes the job automatically. You always get a working result — never a dead end.

v.

The Memory

Every model output lands in a system that knows your business. So a small free model with good context beats a big model flying blind.

Swap inany model Run the buildone prompt, one shot Render-checkdid it actually show? Keep or dropship only what works a model only earns a spot if it renders — the engine never takes the hype's word for it
The gauntlet every model runs: swap in → build → render-check → keep or drop. The trendy 12B got to step three and stopped.
✦
II · the test setup

I grabbed the model everyone's talking about.

Here's what I ran, plain.

The trending model on Hugging Face right now is a 12-billion-parameter coder called gemma-4-12B-coder, "distilled from Fable 5" — Anthropic's top builder. The pitch writes itself: Fable 5 quality, on your own laptop, for free.

So I tested whether that pitch is real.

First surprise: the fast Apple version wouldn't load. Its architecture is too new for the Mac's MLX runtime — it just errors out. The model that's "ready to run locally" isn't, on the fastest path.

The version that does run is the GGUF, through Ollama. That one loaded fine and ran at about 37 tokens a second on my M4 Max — genuinely quick for a 12B, twice the speed of the bigger model I normally use. 7.4GB on disk. Free. Nothing leaves the Mac.

So far, so good. A fast, free, light local model. Now the only question that matters: can it actually build something?

Thinking it? "I'm not technical — I could never set this up."

If you can install one app and paste one line, you can run this. Ollama is a free download. Pulling the model is one command. The whole thing took me longer to type up than to run. And inside the Boardroom there's a step-by-step walk-through for the exact local setup — members who'd never opened a terminal are running models on their own machine.

the trendy 12B vs the proven 27B · fast and small isn't the same as working
Speed — gemma-4-12B37 tok/s
Speed — Qwable 27B18 tok/s
Size on disk — gemma-4-12B7.4 GB
Size on disk — Qwable 27B14 GB
Real scene rendered — gemma-4-12B0%
Real scene rendered — Qwable 27B100%
The 12B wins on speed and size — twice as fast, half the disk. Then the only row that ships a product flips it: it rendered nothing, and the 27B rendered everything.
✦
III · the result

It blanked. So the engine made it build.

My flagship test in GoldieBench is "The Dragon Realm."

One prompt: build a Skyrim-style frozen open world you can walk into, with snow and a sword you draw. One shot. The exact same prompt every model gets. It's the hardest build I run.

Raw, gemma-4-12B handed me a black screen — four times in a row. A broken three.js link one run, a variable scoped wrong another, a zero-by-zero canvas a third. Fast, free, confident, and blank.

That's where most "this model is bad" tests stop. Mine doesn't.

The Local AI Engine did its job: render-check the build, catch the exact bug the model couldn't see, drive the fix, render-check again — until it ran. Here's the same model's build, made to work:

The gemma-4-12B Dragon Realm after the engine fixed it — a walkable frozen world with snowy ground, snow-capped mountains, pine trees and a sword in view
The Dragon Realm the Local AI Engine ships — sparked by the free local model, driven from a black screen to this. Walkable, atmospheric, , offline. WASD walks, mouse looks. Play it live →

Snowy hills out to fogged, snow-capped peaks. A forest of pines. Falling snow. A sword in your hand. You walk with WASD and look with the mouse.

Raw, the free local 12B couldn't get past a black screen. The Local AI Engine took it from there — caught the bug the model couldn't see, drove the fix, and built it out into the atmospheric, walkable world above. The model was the spark; the system did the work. That's the difference between a model and a system — and it all ran , offline, on my own Mac.

the Dragon Realm · how much of a real, walkable scene actually rendered
gemma-4-12B · raw one-shotblack screen
gemma-4-12B · + the enginewalkable world ✓
Qwable 27B · the proven onefull world, first try ✓
Same free 12B, same prompt. Raw it rendered a black screen; once the engine caught the bug and drove the fix, it shipped a walkable world. The proven 27B nailed it first try. (Base Gemma-4 and a 9B local model blanked raw too — same fix turns them around.) Run on my M4 Max.
✦
IV · raw vs the engine

Raw, it blanked. Then the engine made it ship.

Here's the part most "this local model is bad" videos never show you.

One shot, straight out of the box, the trendy 12B handed me a black screen. So I did the thing the Local AI Engine is built for: I let it iterate.

Render-check the build. Catch the exact bug — a broken three.js link, a variable scoped wrong. Drive the fix. Render-check again. Until it actually ran.

Same model. Same prompt. Here's what the engine got it to ship — next to the proven 27B I already run.

Qwable 27B — a full Skyrim-style frozen open world with snow terrain, low-poly trees, falling snow, mountains and a full HUD
✓ shipped first try
Qwable 27B · the proven one
The bigger model I already run nailed it one-shot — a full frozen open world. Some models need the engine to push them; the proven ones just go.

That's the whole thesis in one before-and-after.

The trendy 12B isn't magic — "Fable 5 distilled" made zero difference one-shot. But it isn't useless either. On its own it blanks. Inside a system that render-checks and drives the fix, the exact same free, local 12B ships a real, walkable game.

(Base Gemma-4 and a 9B local model blanked on the raw one-shot too — and the same render-check-and-fix loop is exactly how you turn any of them into something that ships.)

Thinking it? "So the model can't actually do it — the system did."

Right — and that's the entire point. No single model is the moat; they all stumble, even the frontier ones. The moat is the system around them: the render-check that catches the black screen, the loop that feeds the model its own error, the fix that drives it to a working build. The model is this week's fuel. The engine is what turns fuel into something that runs.

✦
V · what I did with it

I still wired it into my Agent OS — on my terms.

A blank screen on the first try doesn't make a model useless. It makes it a model you don't ship blind.

So gemma-4-12B is wired straight into my Local AI Engine as a swappable model — it shows up in the dashboard, it runs live, , offline.

And when I hand it a build, the engine doesn't just hope. It render-checks what came back, and if it's a black screen, it drives the model to the fix — the loop you just watched turn a blank into a walkable world.

That's the whole point. I don't pick one perfect model and pray. I run whatever's free and local behind one system, and let the engine push it to something that actually works — with a render-check standing guard so a broken build never reaches anything that matters.

Thinking it? "Doesn't running an Agent OS burn a fortune in tokens?"

No — that's the biggest myth about it. The Local AI Engine runs the everyday 90% on a free local model on your own machine (nothing leaving it), free APIs and cheap open models slot in for more, and for the frontier work it drives the CLIs you already pay for — your Claude subscription already includes the Claude CLI, and Agent OS plugs straight into it, so you're not paying twice. It's a layer on top of what you already own, not a new meter. And inside the AI Profit Boardroom there are full token-optimisation tutorials, so you cut usage to the bone and never think about it again.

✦
VI · get the engine

Run any model on your own Mac — with the system that tests it.

Everything I just used to test that model — the local engine, the render-check, the swap-any-model dashboard, the fallback to a proven model — is the Agent OS inside the AI Profit Boardroom. Here's what you get:

The Local AI Engine — run any model offline on your own Mac, , nothing leaving it
The build test setup — drop in any model, get a real render-check, not a benchmark number
Agent Kanban — Planner → Builder → Reviewer, agents that catch the bad output before you ship it
The Claude Workspace — every build saved and previewed, nothing lost
Every CLI you already pay for — Claude, Codex, Gemini, Kimi, GLM, Grok, wired into one dashboard
Free local + free API models — the everyday 90% at $0, swap in whatever trends next
The AI Mastermind — models debate, you get the consensus answer
Memory that knows your business — so a small free model with context beats a big one flying blind
Token-efficiency playbooks — cut usage to the bone, stop worrying about cost
3,900+ founders + me, daily — every new model added the week it ships

You're not buying a tool. You're getting the whole operating system I run a seven-figure business on — the one that tested this model in minutes instead of an evening.

Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
✦
VII · old way vs new way

Two ways to handle the model of the week.

Chasing the model
~ an evening lost
  • See the trending model, feel behind
  • Download it, fight the setup for an hour
  • Point it at a real job, get a black screen
  • Can't tell if it's the model or your prompt
  • Ship the broken build, or scrap the night
  • A newer model drops — start over
Running the engine
~ 5 minutes
  • New model drops? Swap it into one slot
  • Run the same real build test, one shot
  • Get a screenshot — did it render or not
  • Keep it if it ships, drop it if it blanks
  • Bad builds get caught before any client sees them
  • A proven model covers the jobs the new one can't
✦
VIII · three beliefs to drop

What's holding you back from running your own models.

Wrong: "The newest trending model is the one I need."

Right: The newest model is a guess until it's tested. This one trended at the top of Hugging Face and rendered a black screen. Trending tells you what's loud, not what works. A render-check tells you the truth.

Wrong: "If it's 'distilled from a frontier model,' it's frontier quality."

Right: "Distilled from Fable 5" sounds amazing and meant nothing here — the base 12B and the "Fable 5" 12B both blanked. A small model is a small model. The label on the box doesn't build the game.

Wrong: "I need the best model before I can do anything."

Right: You need a system, not a model. A free local model inside a system that tests, catches, and falls back beats a frontier model you're using blind. The engine is the part that compounds — the model is just this week's fuel.

Don't take my word for it

158 pages of members who stopped chasing tools and started building systems — real businesses, real wins, documented in their own words.

Read the 158-page testimonials doc →
Trendy 12B · blanked Proven 27B · ships Free local · $0 models — swappable fuel YOUR LOCAL AI ENGINE test · catch · fall back · the part you own whichever model wins this week, it plugs in here
Models are fuel — trendy, proven, free, swappable. The engine that tests them and routes the work is the part nobody can hype away from you.
✦
IX · should you download it?

Run it — just don't trust it.

Should you grab gemma-4-12B? Sure. It's free, it's fast, it's light. For quick answers and simple drafts it's perfectly fine. Drop it in your engine and let it earn its keep on the easy stuff.

Just don't hand it the work that matters until it's passed a real test. That goes for every model — this one, the next one, and the one trending next week.

The people who figure out local AI now, while the models change every week, are going to be way ahead when things settle. Every model you test. Every render-check you run. Every job you route to the right model. It all compounds — into a system, not a pile of downloads.

Most people will keep chasing the model of the week. The people running an engine test once and move on.

✦
X · recap

What you walk away with.

i.

You stopped trusting hype. The #1 trending local model rendered a black screen. Trending ≠ working.

ii.

You test in minutes. Same prompt, one shot, a screenshot for an answer — not a lost evening.

iii.

You stopped paying. 37 tok/s, 7.4GB, $0 — the model runs free on your own Mac.

iv.

You catch the black screens. A render-check stands guard so broken builds never reach a client.

v.

You never depend on one model. Any model swaps in; a proven one covers what the new one can't.

vi.

You own the engine. Models are this week's fuel. The system is the part that compounds.

The model isn't the moat. The engine that tests it is.

XI · your move

Stop chasing models. Build the engine once.

You just watched a trendy model blank on the build that matters — and a tested one ship a whole world. The difference wasn't the model. It was the system around it. That system is the Agent OS, and you can run the whole thing on your own Mac.

Inside the AI Profit Boardroom you get the Local AI Engine that runs any model offline for free, the build-test setup that tells you in minutes whether a model actually works, Agent Kanban that catches the bad output, the Claude Workspace that saves every build, every CLI you already pay for wired into one dashboard, memory that knows your business, full token-efficiency playbooks, and 3,900+ founders building alongside you — with every new model added the week it drops. One system. It runs on whatever wins next.

The people chasing the model of the week are still downloading. You could be the one who tested it, kept what works, and shipped.

Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab

Link's in the description. Test the next model before you trust it — I'll see you in the next one.