Grok 4.6 just tied the best AI models on earth while costing a fraction of the price.
I handed it sixteen video games and gave it one shot at each one.
No fixing, no second tries, no help from me.
What came back is on this page, and you can play every single one of them yourself.
Two of them broke in a way I have never seen a model recover from.
Stick with me to the end for that part, because it changes how you should think about every model you use.
What you're watching: real gameplay footage from eight of the sixteen games Grok 4.6 wrote in one shot. Every frame on this page came out of a build it authored itself.
Artificial Analysis is one of the independent groups that tests every model on the same nine benchmarks. Grok 4.5 scored 56 a month ago. Grok 4.6 scored 61.
What you're looking at: the same nine-benchmark index every model on the chart is measured with. Grok 4.6 lands level with GPT-5.6 Sol and one point under Claude Fable 5 Max.
This is the announcement post from the lab, on the day it shipped.
Worth reading first so everything below is measured against what they actually claimed, not what the internet said they claimed.
Elon's claim is that when you weigh intelligence, speed and cost together, nothing beats it. He owns the company, so treat that as marketing — but the cost half of it holds up in the independent numbers.
This is the claim itself, straight from Elon.
Keep it in mind when you get to the games section, because that's where I stopped taking anyone's word for it.
That's per million tokens — roughly half what the other frontier models charge for the same index score. All sixteen games on this page cost me less than a takeaway coffee.
What you're looking at: output pricing, the number that actually decides your monthly bill, because output is what agents generate all day.
Each of these games is a single HTML file of thirty to forty-five thousand characters. It wrote most of them in three to five minutes, start to finish.
What you're looking at: the whole loop for every game on this page. One prompt in, one finished file out, no conversation in between.
Same day it dropped, I gave Grok 4.6 sixteen game briefs from my own benchmark and took the first thing it handed back — then played every single one through its full gameplay arc before scoring it.
What you're looking at: the exact process behind every score below. A game gets driven the way a player drives it — take off, crash, deploy the chute — and the frame you see is from the middle of that play.
Real footage, captured while the games were being played. Nothing here is a mockup and nothing was touched by hand.
What you're watching: the identical one-shot brief handed to Grok 4.6, GPT-5.6 Sol and Claude Fable 5, all three played side by side. Grok's has light-gates and traffic to dodge; GPT's is the cleanest road; Fable 5 goes widest on colour.
What you're watching: Grok's is the moodiest and by far the darkest, GPT's corridor is the brightest and easiest to read, and Fable 5 puts a full-screen red DOOMED wash over yours. Same one-line brief for all three.
What you're watching: three takes on 'steal cars, outrun cops'. Watch the speed readouts and the minimaps — all three wired a real driving model, and they still look nothing like each other.
What you're watching: same fantasy brief, three different worlds. Look at the hero models — this is the test that separates a real multi-part character from a coloured capsule.
What you're watching: enemy waves, radar and hull damage in all three. The differences are in the juice — trails, tracers and how much the screen reacts when you hit something.
This one puts Grok 4.6 next to GPT-5.6 Sol on a platformer, played side by side, exactly the way the videos above are set up.
Different game, different tester, same method — because watching two builds run at the same time tells you more than any score does.
Every card opens the real build on my benchmark site. Scores are out of ten.
Outrunsynthwave racer
6.4
Doomunplayable until it rebalanced itself
8.0
Neon Blasterarcade space shooter
7.6
City Drivingopen-world sandbox
7.4
Street Levelthird-person shooter
7.4
Twilight Valeaction RPG
6.6
Dragon Realmfrozen open world
6.4
Ember Valetop-down RPG
6.0
The Cryptdungeon crawler
5.4
Parachute Dropchute was dead on the first try
5.4
Vector Rushno track under the car
5.4
Neon Citycrashes into everything
5.0
Flight Simgorgeous world, dead flight model
4.8
Dogfightspawned into the ground
1.5
Voxel Craftblack screen on the first try
1.5
Skyrim-styleone typo killed the whole file
What you're looking at: all sixteen, winners and wrecks. Twelve of the sixteen were playable on the first try. The last two never rendered at all — and the wrecks are the interesting part.
A game is the hardest possible version of "follow my instructions exactly."
It has to hold a spec across forty thousand characters, wire every control it advertises, and still be running a minute later — or you see the failure instantly.
A model that can do that will happily hold a spec across your report, your client email or your data cleanup.
Five builds needed a repair pass. I sent each one back with nothing but the real evidence — the console error, the HUD numbers that never moved, how fast I died — and asked for a corrected file. It fixed every one of them.
What you're looking at: the loop that turned five broken builds into five working games. The only thing I contributed was the evidence.
What you're watching: the build that scored 1.5. First try it rendered a pure black screen with one console error. I sent it that single line of text. This is what came back — a lit voxel world with a miner, a block hotbar and a day cycle.
What you're looking at: the honest scoreboard. The score on the leaderboard is still the first try — the repairs are what you get to play.
Databricks ran it on their document-and-data benchmark — reading reports, pulling numbers out of messy files — and it set the top score. That's closer to your Tuesday than any game is.
A smart model you can't reach is useless. Here's every way in as of today.
Grok Build — xAI's own build surface, with usage limits. x.ai/build →
Cursor — half price for the launch week, 256,000-token window. cursor.com →
Grok Bot — their agent product, newer and still growing.
The API — this is the door that lets it become the brain inside the setup you already run. console.x.ai →
What you're looking at: the four ways in. Everything I built on this page came through the last one.
The model is the brain. Your agent setup is the body. You just got a cheaper, faster brain — and you don't have to rebuild the body to use it.
What you're looking at: why the model release matters less than the socket you plug it into. Build the socket once and every launch is a free upgrade.
One place where your agents live, so a model is a setting and not a rebuild. This is the part you only build once.
Send the everyday jobs to the cheap fast brain and keep the expensive one for the hard thinking. Same work, a fraction of the spend.
When something breaks, hand the model the real error instead of fixing it yourself. That's what turned four dead games into four working ones.
No — that's the biggest myth about it. The everyday 90% runs on a free local model on your own machine, at zero cost, with nothing leaving it.
Free APIs slot in on top, and for frontier work it drives the CLIs you already pay for — your Claude subscription already includes the Claude CLI, so you're not paying twice.
And inside the Boardroom there are full token-efficiency tutorials, so you cut usage to the bone and stop thinking about it.
You can wire this together yourself with everything on this page. Or get the whole thing done inside the Agent Operating System, with Grok 4.6 already sitting in the socket.
What you're watching: the third-person street brief. Ammo counters, pedestrians and a working minimap in all three — the gap is in the lighting and how alive the street feels.
xAI used Grok 4.5 to generate huge amounts of training work, then trained 4.6 on agentic reinforcement learning — coding, knowledge work, multi-step tasks where the model has to use tools and check its own work.
What you're looking at: the difference that shows up in a forty-thousand-character game file. Holding a spec that long is the whole skill.
Simple job, keep it low and fast. Big messy problem, crank it up. You're choosing how much brain power you pay for on every single task.
What you're looking at: the dial you control. Every game on this page was built without touching extra high — that's still headroom I haven't spent.
Every game brief demanded the same list: a multi-part hero model, layered world, fog for depth, working controls, one clean HUD, a fail state. Look at the footage and tick them off — the spec is visibly there.
What you're watching: the knight has a cape and a drawn sword instead of being a grey capsule, the forest has near, mid and far layers, and the health bar drains because the crystals actually reach you. That's four spec lines you can see at once.
Then count them on the screen. Twelve of sixteen builds shipped every element of the brief on the first pass.
The complaint about AI has always been that you say one thing and it does another. A model that does what you asked at a tenth of the price is worth more than a slightly smarter one that ignores half your prompt.
What you're watching: freefall and canopy, three ways. Grok's canopy only opens because it repaired its own deploy key — that's the fix from earlier, running live.
The games are just proof it holds a spec. The value is the boring, high-volume jobs you were rationing because of cost.
Your research agent and your report agent just got dramatically cheaper to leave running all day.
Run client work through a frontier model without watching the meter every prompt.
Product descriptions and customer replies on top-tier intelligence at bottom-tier cost.
The always-on workload that used to need funding — research all day, content all night.
What you're looking at: the four shapes this actually takes. None of them require you to write a line of code.
A harness is just the app you use the model through. ChatGPT has a great desktop app, Claude has a great one, and Grok's options are more scattered — Cursor is excellent but coding-first, and Grok Bot is newer.
What you're looking at: the reason the harness gap is an argument for owning your own setup, not against using the model.
What you're watching: the build that started as a black screen. After one round of its own error message, Grok's sits next to the other two as a real world you can walk around.
No plugins, no workarounds. If your business depends on knowing what people are saying right now — trends, complaints, what's working in your market — it pulls that live. Claude can't. GPT can't do it natively.
It can generate images and turn them into video. One creator wired an agent to write the prompt, make the images, animate them, cut them together with music and post — set up once, running for weeks untouched.
What you're watching: the same principle, one level down. I described a game, walked away, and came back to this. The daily video loop is that idea pointed at content instead of code.
4.5 to 4.6 took about a month and jumped five points. If 4.7 lands anywhere near that, xAI goes from catching up to leading — and they have the hardware to keep the pace.
What you're looking at: the pace. This is why the socket matters more than the model sitting in it today.
I understand the feeling. But look at what happened this week — the people who were already set up plugged Grok 4.6 in on day one. The people waiting got nothing, because they had nothing to plug it into.
Wrong: "I'll wait until the models settle down, then learn the winner."
Right: They won't settle. A new one lands every month. Learn the setup once and every launch becomes a free upgrade instead of a fresh start.
Wrong: "Keeping up with models is the skill."
Right: Keeping up with models is a losing game. The skill is having a body ready for whatever brain drops next.
Wrong: "I'll pick one model and stick with it, like a football team."
Right: The businesses winning right now mix and match. Heavy strategy on the expensive model, high-volume everyday work on the cheap fast one.
158 pages of members who already broke through these exact beliefs. Real businesses, real numbers, in their own words.
Read the 158-page testimonials doc →Everything on this page started as plain English typed into a box. The office-work benchmark it topped was about reading documents, not writing code.
If you can write an email, you can run this. You describe what you want and the agent does the work.
Members who had never opened a terminal are running the same setup — and the coaching calls exist for exactly the moment you get stuck.
Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.
Read the 158-page wins doc →Fair — and for a lot of jobs ChatGPT is still great. But you're paying frontier prices for every single task, including the simple ones.
What you're looking at: the actual change. Not a new tool — a different way of spending on the tools you already have.
For two years, if you wanted the best you paid the most. That link just broke — and when the price of intelligence drops this hard, the winners aren't the biggest budgets any more, they're the best systems.
Always-on AI workloads needed funding. Research all day, content all night, follow-up that never sleeps — that was a company with a budget.
A solo operator with a good agent setup runs the same load, powered by a model that ties the best on earth.
You need somewhere to plug it in. That's the only thing standing between you and this.
What you're looking at: the window that opened this week — and with 4.7 weeks away, it's only opening wider.
This page shows you what Grok 4.6 can do. The Boardroom saves you the year it took me to build everything around it — the socket, the routing, the memory, the repair loop.
Readers bookmark this page and carry on paying frontier prices for simple jobs. Operators join, install the Agent OS this week, and route their everyday work to a brain that costs a fraction.
The AI race just stopped being a two-horse race. The best intelligence in the world just got cheap. The only question left is whether you have somewhere to plug it in.