Grok 4.6 · shipped 12 Aug 2026

The One-Tenth Engine + Grok 4.6.

Grok 4.6 just tied the best AI models on earth while costing a fraction of the price.

I handed it sixteen video games and gave it one shot at each one.

No fixing, no second tries, no help from me.

What came back is on this page, and you can play every single one of them yourself.

Two of them broke in a way I have never seen a model recover from.

Stick with me to the end for that part, because it changes how you should think about every model you use.

What you're watching: real gameplay footage from eight of the sixteen games Grok 4.6 wrote in one shot. Every frame on this page came out of a build it authored itself.

0intelligence index
$2 / $6per million tokens
0games, one shot each
0K token context
§1I ────── what actually happened

Five points in one month.

Artificial Analysis is one of the independent groups that tests every model on the same nine benchmarks. Grok 4.5 scored 56 a month ago. Grok 4.6 scored 61.

ARTIFICIAL ANALYSIS INTELLIGENCE INDEX Grok 4.5 · last month 56 Grok 4.6 · today 61 GPT-5.6 Sol 61 Claude Fable 5 Max 62 MOST LABS TAKE 3–6 MONTHS xAI DID IT IN ONE +5 points in about four weeks

What you're looking at: the same nine-benchmark index every model on the chart is measured with. Grok 4.6 lands level with GPT-5.6 Sol and one point under Claude Fable 5 Max.

Straight from the source — read it and run it yourself ↓
Tweet 1 · the launch

xAI posts the release itself

This is the announcement post from the lab, on the day it shipped.

Worth reading first so everything below is measured against what they actually claimed, not what the internet said they claimed.

Tweet 2 · the numbers

The benchmark card

The follow-up post carries the benchmark table for the release.

The scores in my chart above come from the independent testers, not this card, so you can check one against the other.

§2II ────── the claim

He says it's number one. He also owns it.

Elon's claim is that when you weigh intelligence, speed and cost together, nothing beats it. He owns the company, so treat that as marketing — but the cost half of it holds up in the independent numbers.

Tweet 3 · the founder

The "objectively number one" post

This is the claim itself, straight from Elon.

Keep it in mind when you get to the games section, because that's where I stopped taking anyone's word for it.

Tweet 4 · the reaction

What the builders said within hours

Here's the response from people who had already put it to work.

The pattern in the replies is the same one I hit: it does what you asked, and the bill barely moves.

§3III ────── the price

Two dollars in. Six dollars out.

That's per million tokens — roughly half what the other frontier models charge for the same index score. All sixteen games on this page cost me less than a takeaway coffee.

PRICE PER MILLION OUTPUT TOKENS · SAME INDEX SCORE $6 Grok 4.6 roughly 2× the usual frontier price up to 10× on the most expensive tiers what testers hit

What you're looking at: output pricing, the number that actually decides your monthly bill, because output is what agents generate all day.

§4IV ────── the speed

Around 100 tokens a second.

Each of these games is a single HTML file of thirty to forty-five thousand characters. It wrote most of them in three to five minutes, start to finish.

one prompt the game spec 3–5 minutes no back and forth a playable game one file, ~40KB

What you're looking at: the whole loop for every game on this page. One prompt in, one finished file out, no conversation in between.

A benchmark you can't play is just a number on a slide.
§5V ────── the test I ran

Sixteen games. One shot each.

Same day it dropped, I gave Grok 4.6 sixteen game briefs from my own benchmark and took the first thing it handed back — then played every single one through its full gameplay arc before scoring it.

the brief one game, one prompt Grok 4.6 writes it 100% of the code I play it full gameplay arc then a score from a mid-play frame NOTHING IS SCORED FROM A POSED SCREENSHOT

What you're looking at: the exact process behind every score below. A game gets driven the way a player drives it — take off, crash, deploy the chute — and the frame you see is from the middle of that play.

Tweet 5 · the leaderboard

The independent boards moved the same day

This is one of the public arenas reacting to the release.

My bench is a different shape — it only cares whether the thing it builds actually runs — but the direction matched.

§6VI ────── the games

This is what came back. First try.

Real footage, captured while the games were being played. Nothing here is a mockup and nothing was touched by hand.

same brief · three models · played at the same time

Outrun, built by all three.

What you're watching: the identical one-shot brief handed to Grok 4.6, GPT-5.6 Sol and Claude Fable 5, all three played side by side. Grok's has light-gates and traffic to dodge; GPT's is the cleanest road; Fable 5 goes widest on colour.

same brief · three models

Doom, built by all three.

What you're watching: Grok's is the moodiest and by far the darkest, GPT's corridor is the brightest and easiest to read, and Fable 5 puts a full-screen red DOOMED wash over yours. Same one-line brief for all three.

same brief · three models

Open-city driving, built by all three.

What you're watching: three takes on 'steal cars, outrun cops'. Watch the speed readouts and the minimaps — all three wired a real driving model, and they still look nothing like each other.

same brief · three models

The twilight world, built by all three.

What you're watching: same fantasy brief, three different worlds. Look at the hero models — this is the test that separates a real multi-part character from a coloured capsule.

same brief · three models

The arcade shooter, built by all three.

What you're watching: enemy waves, radar and hull damage in all three. The differences are in the juice — trails, tracers and how much the screen reacts when you hit something.

Tweet 9 · other people are doing this too

The same head-to-head, run by someone else

This one puts Grok 4.6 next to GPT-5.6 Sol on a platformer, played side by side, exactly the way the videos above are set up.

Different game, different tester, same method — because watching two builds run at the same time tells you more than any score does.

Play all sixteen yourself.

Every card opens the real build on my benchmark site. Scores are out of ten.

8.6Grok 4.6 one-shot build of Outrun — GoldieBench AI benchmark screenshotOutrunsynthwave racer 6.4Grok 4.6 one-shot build of Doom — GoldieBench AI benchmark screenshotDoomunplayable until it rebalanced itself 8.0Grok 4.6 one-shot build of Neon Blaster — GoldieBench AI benchmark screenshotNeon Blasterarcade space shooter 7.6Grok 4.6 one-shot build of an open-city driving sandbox — GoldieBench AI benchmark screenshotCity Drivingopen-world sandbox 7.4Grok 4.6 one-shot build of a third-person street shooter — GoldieBench AI benchmark screenshotStreet Levelthird-person shooter 7.4Grok 4.6 one-shot build of Twilight Vale — GoldieBench AI benchmark screenshotTwilight Valeaction RPG 6.6Grok 4.6 one-shot build of The Dragon Realm — GoldieBench AI benchmark screenshotDragon Realmfrozen open world 6.4Grok 4.6 one-shot build of a top-down RPG — GoldieBench AI benchmark screenshotEmber Valetop-down RPG 6.0Grok 4.6 one-shot build of a torch-lit crypt crawler — GoldieBench AI benchmark screenshotThe Cryptdungeon crawler 5.4Grok 4.6 build of Parachute Drop with the canopy open — GoldieBench AI benchmark screenshotParachute Dropchute was dead on the first try 5.4Grok 4.6 one-shot build of a 3D racer — GoldieBench AI benchmark screenshotVector Rushno track under the car 5.4Grok 4.6 one-shot build of Neon City — GoldieBench AI benchmark screenshotNeon Citycrashes into everything 5.0Grok 4.6 build of a flight simulator on the runway — GoldieBench AI benchmark screenshotFlight Simgorgeous world, dead flight model 4.8Grok 4.6 build of an air-combat dogfight game — GoldieBench AI benchmark screenshotDogfightspawned into the ground 1.5Grok 4.6 build of a voxel mining game — GoldieBench AI benchmark screenshotVoxel Craftblack screen on the first try 1.5Grok 4.6 build of a Skyrim-style crawler that failed to render — GoldieBench AI benchmark screenshotSkyrim-styleone typo killed the whole file

What you're looking at: all sixteen, winners and wrecks. Twelve of the sixteen were playable on the first try. The last two never rendered at all — and the wrecks are the interesting part.

Thinking it? "Games are a party trick. What does that prove about my business?"

A game is the hardest possible version of "follow my instructions exactly."

It has to hold a spec across forty thousand characters, wire every control it advertises, and still be running a minute later — or you see the failure instantly.

A model that can do that will happily hold a spec across your report, your client email or your data cleanup.

The wrecks taught me more than the wins.
§7VII ────── the thing almost no model does

I handed it its own error. It fixed itself.

Five builds needed a repair pass. I sent each one back with nothing but the real evidence — the console error, the HUD numbers that never moved, how fast I died — and asked for a corrected file. It fixed every one of them.

it breaks black screen, dead key I capture proof the real error line same model reads it no human edits it works 4 out of 4, first retry THE MODEL WRITES EVERY FIX — I NEVER TOUCH THE FILE

What you're looking at: the loop that turned five broken builds into five working games. The only thing I contributed was the evidence.

Voxel Craft · black screen → this, on one retry

What you're watching: the build that scored 1.5. First try it rendered a pure black screen with one console error. I sent it that single line of text. This is what came back — a lit voxel world with a miner, a block hotbar and a day cycle.

What it got wrong 5 of 16
  • Voxel world: one bad function call, whole screen black
  • Flight sim: gorgeous valley, plane never left the runway
  • Dogfight: spawned into the ground, shot down instantly
  • Parachute: the deploy key did nothing for 30 seconds
  • Doom: the demons killed you in under a second, every time
  • Skyrim-style: one duplicate variable killed the file
After it read its own error 5 fixed
  • Voxel world renders, you can mine and place blocks
  • Flight sim rotates and climbs to 992 feet
  • Dogfight spawns airborne and holds full hull
  • Parachute canopy opens, descent drops 52 to 5 m/s
  • Doom now survives a full fight — you kill demons instead
  • Every fix written by the model, not by me

What you're looking at: the honest scoreboard. The score on the leaderboard is still the first try — the repairs are what you get to play.

Tweet 6 · the hands-on take

Someone who put it straight to work

Here's a builder's write-up from the first day.

The theme that keeps coming up is instruction-following, which is exactly what the self-repair depends on.

§8VIII ────── the boring benchmark that matters more

It also topped a test about office work.

Databricks ran it on their document-and-data benchmark — reading reports, pulling numbers out of messy files — and it set the top score. That's closer to your Tuesday than any game is.

Tweet 7 · the office-work result

The document benchmark

This is the post covering that result.

Games prove it can follow a hard spec. This proves it can read your files without inventing things.

§9IX ────── where you can actually use it

Four doors. One of them matters most.

A smart model you can't reach is useless. Here's every way in as of today.

free to try

Grok Build — xAI's own build surface, with usage limits. x.ai/build →

for coding

Cursor — half price for the launch week, 256,000-token window. cursor.com →

agent app

Grok Bot — their agent product, newer and still growing.

the one that matters

The API — this is the door that lets it become the brain inside the setup you already run. console.x.ai →

What you're looking at: the four ways in. Everything I built on this page came through the last one.

§10X ────── the framework

The One-Tenth Engine.

The model is the brain. Your agent setup is the body. You just got a cheaper, faster brain — and you don't have to rebuild the body to use it.

ONE BODY · INTERCHANGEABLE BRAINS Grok 4.6 Claude whatever ships next THE SOCKET your agent setup memory · tools · jobs research that runs all day content and drafts lead follow-up

What you're looking at: why the model release matters less than the socket you plug it into. Build the socket once and every launch is a free upgrade.

i.

The Socket

One place where your agents live, so a model is a setting and not a rebuild. This is the part you only build once.

ii.

The Route

Send the everyday jobs to the cheap fast brain and keep the expensive one for the hard thinking. Same work, a fraction of the spend.

iii.

The Loop

When something breaks, hand the model the real error instead of fixing it yourself. That's what turned four dead games into four working ones.

Thinking it? "Doesn't running an Agent OS burn a fortune in tokens?"

No — that's the biggest myth about it. The everyday 90% runs on a free local model on your own machine, at zero cost, with nothing leaving it.

Free APIs slot in on top, and for frontier work it drives the CLIs you already pay for — your Claude subscription already includes the Claude CLI, so you're not paying twice.

And inside the Boardroom there are full token-efficiency tutorials, so you cut usage to the bone and stop thinking about it.

§11XI ────── skip the setup Skip the setup

Get the One-Tenth Engine built for you.

You can wire this together yourself with everything on this page. Or get the whole thing done inside the Agent Operating System, with Grok 4.6 already sitting in the socket.

Grok 4.6 pre-wired as a swap-in brain — into Hermes, OpenClaw and the Agent OS, configured like my machine
The full Agent OS zip — installs in an afternoon, with a 30-day roadmap
Every CLI you already pay for, in one dashboard — Claude, Codex, Gemini, Kimi, GLM, Grok
Free local models for the everyday 90% — so most of your work costs nothing at all
Four coaching calls a week — ask about your own Grok setup, live, and get unstuck the same day
Daily tutorials and a prompt library — new tools added the week they ship
3,900+ founders across 38 countries — someone's always online who's tested what you're trying
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Set up in an afternoon · used in 38 countries · new tools added the week they ship
between the sections · same brief, three models

On foot, built by all three.

What you're watching: the third-person street brief. Ammo counters, pedestrians and a working minimap in all three — the gap is in the lighting and how alive the street feels.

§12XII ────── what's new under the hood

They trained it to finish jobs, not answer questions.

xAI used Grok 4.5 to generate huge amounts of training work, then trained 4.6 on agentic reinforcement learning — coding, knowledge work, multi-step tasks where the model has to use tools and check its own work.

MEMORISED THE TEXTBOOK answers your question beautifully loses the plot on step nine never checks its own work DID THE WORK EXPERIENCE holds the spec for fifty steps uses tools and reads the result notices when it got it wrong BOTH CAN ANSWER. ONLY ONE CAN DO THE JOB.

What you're looking at: the difference that shows up in a forty-thousand-character game file. Holding a spec that long is the whole skill.

§13XIII ────── the setting worth knowing

Four thinking levels, including a new extra high.

Simple job, keep it low and fast. Big messy problem, crank it up. You're choosing how much brain power you pay for on every single task.

LOW quick replies, cheapest MEDIUM normal day-to-day work HIGH where 4.5 topped out EXTRA HIGH brand new · the hardest jobs CHEAPER + FASTER SLOWER + SMARTER

What you're looking at: the dial you control. Every game on this page was built without touching extra high — that's still headroom I haven't spent.

§14XIV ────── the boring superpower

It does what you asked.

Every game brief demanded the same list: a multi-part hero model, layered world, fog for depth, working controls, one clean HUD, a fail state. Look at the footage and tick them off — the spec is visibly there.

The same spec, obeyed: hero model · layered world · live HUD · enemies that fight back

What you're watching: the knight has a cape and a drawn sword instead of being a grey capsule, the forest has near, mid and far layers, and the health bar drains because the crystals actually reach you. That's four spec lines you can see at once.

Thinking it? "Every model claims it follows instructions."

Then count them on the screen. Twelve of sixteen builds shipped every element of the brief on the first pass.

The complaint about AI has always been that you say one thing and it does another. A model that does what you asked at a tenth of the price is worth more than a slightly smarter one that ignores half your prompt.

between the sections · same brief, three models

The parachute jump, built by all three.

What you're watching: freefall and canopy, three ways. Grok's canopy only opens because it repaired its own deploy key — that's the fix from earlier, running live.

§15XV ────── what this means for your day

Same work. A fraction of the bill.

The games are just proof it holds a spec. The value is the boring, high-volume jobs you were rationing because of cost.

agency

Your research agent and your report agent just got dramatically cheaper to leave running all day.

freelancer

Run client work through a frontier model without watching the meter every prompt.

ecommerce

Product descriptions and customer replies on top-tier intelligence at bottom-tier cost.

solo operator

The always-on workload that used to need funding — research all day, content all night.

What you're looking at: the four shapes this actually takes. None of them require you to write a line of code.

Frontier intelligence stopped being expensive. That's the whole story.
§16XVI ────── the honest weakness

A brilliant model with a harness problem.

A harness is just the app you use the model through. ChatGPT has a great desktop app, Claude has a great one, and Grok's options are more scattered — Cursor is excellent but coding-first, and Grok Bot is newer.

The honest gap no single app
  • No one Grok app that does everything
  • Cursor is superb, but built for code first
  • Grok Bot is good and still growing
  • You'd be hopping between surfaces all day
  • Two builds died on a hard error, three more were unplayable
Why it stops mattering your setup is the app
  • Plug it into the agent setup you already run
  • Your dashboard becomes the harness
  • The model is just the brain inside it
  • Swap it out the day something better lands
  • Hand it the error and it repairs its own work

What you're looking at: the reason the harness gap is an argument for owning your own setup, not against using the model.

between the sections · same brief, three models

The voxel world, built by all three.

What you're watching: the build that started as a black screen. After one round of its own error message, Grok's sits next to the other two as a real world you can walk around.

§17XVII ────── the thing nobody can copy

It's wired into X, live.

No plugins, no workarounds. If your business depends on knowing what people are saying right now — trends, complaints, what's working in your market — it pulls that live. Claude can't. GPT can't do it natively.

Tweet 8 · the wider reaction

The rundown from the day it landed

This one gathers the reaction in one place.

It's also a neat demonstration of the point: this is the conversation Grok can read live and the other models can't.

§18XVIII ────── the loop people are already running

Pictures in, video out, every day, unattended.

It can generate images and turn them into video. One creator wired an agent to write the prompt, make the images, animate them, cut them together with music and post — set up once, running for weeks untouched.

Unattended output: this footage came from a build nobody supervised

What you're watching: the same principle, one level down. I described a game, walked away, and came back to this. The daily video loop is that idea pointed at content instead of code.

§19XIX ────── what's coming

4.7 is weeks away.

4.5 to 4.6 took about a month and jumped five points. If 4.7 lands anywhere near that, xAI goes from catching up to leading — and they have the hardware to keep the pace.

Grok 4.556 · last month Grok 4.661 · today Grok 4.7a few weeks away

What you're looking at: the pace. This is why the socket matters more than the model sitting in it today.

§20XX ────── three beliefs to drop

"It's moving too fast. I'll wait."

I understand the feeling. But look at what happened this week — the people who were already set up plugged Grok 4.6 in on day one. The people waiting got nothing, because they had nothing to plug it into.

Wrong: "I'll wait until the models settle down, then learn the winner."

Right: They won't settle. A new one lands every month. Learn the setup once and every launch becomes a free upgrade instead of a fresh start.

Wrong: "Keeping up with models is the skill."

Right: Keeping up with models is a losing game. The skill is having a body ready for whatever brain drops next.

Wrong: "I'll pick one model and stick with it, like a football team."

Right: The businesses winning right now mix and match. Heavy strategy on the expensive model, high-volume everyday work on the cheap fast one.

Don't take my word for it

158 pages of members who already broke through these exact beliefs. Real businesses, real numbers, in their own words.

Read the 158-page testimonials doc →
§21XXI ────── "but I'm not a coder"

The interface is a sentence.

Everything on this page started as plain English typed into a box. The office-work benchmark it topped was about reading documents, not writing code.

3,900+ founders inside AIPB
258 documented wins
400k YouTube subscribers
38 countries
163k X followers
Thinking it? "This all sounds technical. I'm not a developer."

If you can write an email, you can run this. You describe what you want and the agent does the work.

Members who had never opened a terminal are running the same setup — and the coaching calls exist for exactly the moment you get stuck.

Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.

Read the 158-page wins doc →
§22XXII ────── "I already use ChatGPT"

Old way vs new way.

Fair — and for a lot of jobs ChatGPT is still great. But you're paying frontier prices for every single task, including the simple ones.

Old way ~$6 a task
  • One model, one subscription, every job goes through it
  • Simple drafts cost the same as hard strategy
  • You ration the volume work to protect the bill
  • New model lands, you start learning from scratch
  • Your context lives inside somebody else's app
  • When it breaks, you fix it by hand
New way ~10% of that
  • One socket, many brains, each job routed by difficulty
  • Everyday volume on the cheap fast brain
  • Hard strategy still goes to the expensive one
  • New model lands, you change one setting
  • Your memory and tools stay yours
  • When it breaks, the model reads the error and fixes itself

What you're looking at: the actual change. Not a new tool — a different way of spending on the tools you already have.

§23XXIII ────── so what changed

The best is now also nearly the cheapest.

For two years, if you wanted the best you paid the most. That link just broke — and when the price of intelligence drops this hard, the winners aren't the biggest budgets any more, they're the best systems.

before

Always-on AI workloads needed funding. Research all day, content all night, follow-up that never sleeps — that was a company with a budget.

now

A solo operator with a good agent setup runs the same load, powered by a model that ties the best on earth.

the catch

You need somewhere to plug it in. That's the only thing standing between you and this.

What you're looking at: the window that opened this week — and with 4.7 weeks away, it's only opening wider.

§24XXIV ────── your move Your move

Get the One-Tenth Engine running this week.

This page shows you what Grok 4.6 can do. The Boardroom saves you the year it took me to build everything around it — the socket, the routing, the memory, the repair loop.

Readers bookmark this page and carry on paying frontier prices for simple jobs. Operators join, install the Agent OS this week, and route their everyday work to a brain that costs a fraction.

Grok 4.6 wired in as a swap-in brain — Hermes, OpenClaw and the Agent OS, done for you
The same-day upgrade when 4.7 drops — we walk you through the swap the week it lands
Your video tools, SEO agents and avatars in one dashboard — not five tabs
Four live coaching calls a week — bring your own setup and your own error messages
A prompt library and a member map — find operators near you already running this
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
258 documented member wins · 38 countries · new tools added the week they ship

The AI race just stopped being a two-horse race. The best intelligence in the world just got cheap. The only question left is whether you have somewhere to plug it in.