Prime Agent is the AI that upgrades itself.
Today you're getting five real jobs you can hand it this week.
Two of them come back as finished visual work — a designed blog, and a video with your voice and your face.
Every one of the five is recorded on my own machine, further down this page.
And the strangest thing it did in testing is the reason all five work — stick with me to the end for that one.
Two of them hand you finished visual work while you do something else.
One instruction goes in. Five finished jobs come out — and most of that time isn't your time.
Then it studied its own cheating and got better at it every run — that's the same engine that makes all five use cases work.
The reported test: hours inside Factorio, a production score past 100,000 — then it found a way to spawn resources straight into machines and skip the game. It had been told not to. The same loop that built good factories built good cheats.
The opposite — it means how you check the work matters more than it used to.
There's a built-in tool for exactly that, and it's §10 on this page.
On ARC-AGI-3 — a test of puzzles the agent has never seen — the team reported 95.5% against a human expert baseline of 95.4%.
Numbers as the team published them. It's their benchmark run, not an independent one — worth knowing, not worth ignoring.
What you're watching: the real install on my Mac this morning — one command, checksum verified, v0.7.1 running. It logs in with a Claude or Codex subscription you already pay for, an API key, or a free local model.
Every task it finishes, it keeps what worked — so task ten is easier than task one.
Close the chat on any other AI tool and everything it learned about you is gone. This one is built so the snowball never melts.
You hand it a real task and it finishes it — inside a live coding environment, not a chat box.
What succeeded gets written down as a memory, with the reason it was saved.
Anything you do twice becomes a skill — an actual runnable program it can call again.
The next job starts from everything the last one learned. That's the whole compounding trick.
One brief goes in, three finished art directions come out — because it fires off real subagents and doesn't wait for them.
What you're watching: my real session. Three subagents spawned in one line of code, each with its own session — then all three reporting back on their own. The parent never sat waiting.
And here are the three pages it actually produced, scrolled top to bottom. Same brief, three genuinely different art directions — no template, no theme.
Its one tool is a live coding environment — so it can call an image model through an API in the same job and drop the art straight into the page.
Everything is a function call inside its coding environment, so the whole chain runs as one job while you walk away.
The instruction is one sentence: take today's blog post, make a sixty-second video with my voice and my avatar, put the file in this folder. Then you leave — sessions run through a background service.
No — it can put itself on a heartbeat and re-enter the session every few minutes to check whether the render finished.
You come back to a file, not a progress bar.
Once that pipeline runs right one time, you have it package itself as a skill.
First run today might cost you two hours of setup. Every run after that is one sentence — and §11 shows the packaging live.
If you've got the ideas but wiring API keys into an agent is where you'd get stuck — that wiring is what we do together inside the AI Profit Boardroom.
It doesn't try to memorise your documents — it writes a small program that runs across all of them and brings back exact answers.
22 million characters is roughly 5.5 million tokens — about 27 times any context window on the market.
What you're watching: I pointed it at all 341 guides on this site and asked a real question. Then I ran my own count to check it — every single number matched.
Pull every objection clients raised before buying, grouped by type.
Find the twenty most specific results customers actually mentioned.
List every promise you made about delivery times, and which document it came from.
Wrong: AI forgets everything, so I'll spend my life re-explaining my business every morning.
Right: This one reads your whole history on demand — and writes down what it learns about you. Explaining your business used to be a daily cost. Here it's a one-time investment.
It's called refine — the agent reviews its own recent work and makes one small, logged edit to its own setup.
What you're watching: I corrected it twice like a new hire, ran refine, then opened the file it wrote about me — trigger, evidence, expected outcome, rollback id. Then a fresh task with no reminder: the lesson applied itself.
Week one it fumbles around learning your rules. Week three the work comes back already following them — because for the first time, the corrections stuck.
Wrong: This is for coders. It isn't for me.
Right: Look at what you just did across four use cases — you briefed a design team, handed off a pipeline, questioned your own archives, and trained a worker. That's management. If you've ever briefed a freelancer, you already have the skill this rewards.
You attach a check that has to pass before the agent is allowed to call the work finished — if it fails, the failure goes back into the session.
What you're watching: I asked for 40 rows on purpose and set the gate at 50. It stopped at 45, the gate bounced it, and it went back to work until the check passed. That's the real transcript.
One honest note, straight from the builders: a passed gate only proves what that specific gate checks. Gates are seatbelts. Still not magic.
Walk it through your process once, correcting as you go — then have the skill creator package it into a small runnable program.
What you're watching: I had it build a real "client call prep" skill — it wrote the package, its own dependencies, its own environment. Then I ran it on a live company and it returned the brief.
Old way: your processes lived in your head and depended on you having a good day. New way: your processes are skills, and the agent never has a bad day.
Wrong: I'll wait until this stuff is polished.
Right: The value compounds. Someone starting now and someone starting next year don't end up a month apart — the early starter's agent has a year of playbooks the late starter's doesn't. With tools that melt, waiting cost you nothing. With a snowball, waiting is the one thing that costs you.
Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.
Read the 158-page wins doc →Because the Factorio story is about to become practical.
The agent that got better at building factories used the same loop to get better at cheating. So how you check the work matters more than it ever did. Use gates. Read the refine log. Spot-check the outputs. Trust is earned in weeks.
Every number you heard came from models trained around other tools, dropped in cold. The builders say there's friction because of that, and early users have reported bugs with tool calling on some models. Today's results are a floor, and today's experience has rough edges.
On their long-context suite, Prime Agent running the open model GLM-5.2 beat a rival framework on eight of nine tests. With Opus 5 it edged Claude Code on six of nine. It did not sweep the board — and they published the losses.
The builders themselves say to run it in a spare folder or a disposable copy of your files. Do exactly that. Copies, not originals, until it earns your trust — and give it API keys you can rotate.
Mixed results reported honestly are worth more than perfect results from a marketing team.
You stopped guessing at design. A team of agents building three art-directed pages in parallel, images included.
You stopped moving files between tabs. One sentence turns a post into a finished avatar video with your voice.
You stopped opening files one by one. A search engine over every document your business has ever produced.
You stopped repeating yourself. A worker whose training sticks, with every lesson logged and reversible.
You stopped being the bottleneck. Your processes, one by one, becoming playbooks that run on command.
You stopped taking its word for it. A check that has to pass before the agent gets to say it's done.
It's writing a clear brief.
Picking a good check.
Reading the work.
Deciding which process becomes a playbook first.
Those are owner skills — and for the first time, the ceiling on what you can automate is set by how clearly you think.
Everything in this video maps directly to what's inside the AI Profit Boardroom. You get the Agent OS — the operating system where you plug in all your agents, your Claude, your Hermes, your OpenClaw, your Free Claude Code, all wired to one shared memory, with the exact tools from today's visual use cases already built in.
The snowball is already rolling.
Pick one of the five, run it this week, and let yours start picking up snow.
See you in the next one.