GLM 5.3 plus Hermes Agent is one of the most powerful AI agent setups you can run right now.
Z.ai changed nothing about the model's brain — not one new parameter.
And it still jumped six times higher on terminal tasks, and beat Claude Mythos 5 and GPT-5.6 Sol on a benchmark that shocked people.
It got so good at one thing that Z.ai is holding the open weights back.
There's one number buried in this release that matters more than any score — and almost nobody is talking about it.
Every number on this page comes from Z.ai's own release — nothing second-hand.
"GLM-5.3 runs on the same mixture-of-experts base as GLM-5.2, and every reported gain came from extended post-training rather than a fresh pretrain."
— Z.ai's GLM-5.3 release, 14 August 2026
Same 743 billion parameters as GLM 5.2. Every single gain came from a month of extra training.
The base is untouched. Everything that changed sits in the middle box.
Not quick questions. Long, multi-step work — the kind an experienced engineer needs days to finish.
This is the shape of the work it practised on — long chains, not single answers.
That's exactly the kind of work you want an agent to do.
Not answer one question. Finish a whole project.
That's the exact failure this release went after.
The model was trained inside long jobs, so holding the thread is the thing it practised most.
This tests how well a model works inside a computer terminal on its own. Roughly a six times jump — same brain.
On Z.ai's own coding benchmark, GLM 5.3 hits 34.5% at its highest effort setting. GLM 5.2 got 23.4%.
Agents' Last Exam and DeepSWE both test whether a model can actually run a job, not just answer about one.
CyberGym is a UC Berkeley benchmark: can a model find and confirm real bugs in real source code?
More than a thousand of them were rated critical or high.
They haven't been independently verified yet — and on most other security tests, GLM 5.3 still sits well behind.
Worth saying out loud so nothing here gets overstated.
GLM 5.2 shipped its weights within days. This one didn't — roughly two weeks of safety checks first.
When a lab famous for shipping everything fast decides to slow down, that tells you something.
GLM 5.3 is live right now on the GLM Coding Plan. And that's exactly where the Hermes setup comes in.
What you're watching: the whole migration — one line in the config goes from glm-5.2 to glm-5.3. Skills, keys, MCP servers and memory all carry over untouched.
You can wire this up yourself with the steps below. Or get the finished thing — the Agent OS, with GLM 5.3, Hermes and your other models already connected.
Hermes lets you plug in any model and give it a profile. Here's the real setup on my machine — clone, change one line, run.
What you're watching: the real session — cloning the GLM 5.2 profile, pointing it at glm-5.3, then running a live job that writes a file and runs a shell command.
Because 5.3 shares the exact same base, everything you already had wired up just carries over.
What you're watching: the real Hermes manager with glm-5-3 selected — dropping open to show it sitting beside every other agent profile, with its own skills, keys and config.
No. Same base model means same wiring.
You swap the model name and instantly get the smarter version. Nothing breaks.
GLM 5.3 scores higher than Claude Opus 4.8 on Z.ai's coding benchmark while using well under half the output tokens.
Higher score, top row. Fewer tokens to get there, bottom row. That gap is the whole story.
They write, check, fix, write again — for hours. Every wasted token is wasted time and wasted plan limits.
It's the opposite, and 5.3 made it more true — this runs on a flat coding plan, and the model now uses fewer tokens per task than the old version did.
Inside the Agent OS the everyday work also runs on free local models and the CLIs you already pay for, so you're not paying twice. There are full token-efficiency tutorials in the Boardroom too.
A worker that's cheap to run, doesn't waste tokens, and was literally trained on long multi-day tasks.
Hermes holds the profiles, the skills, the memory and the board. It decides who does what, and when.
GLM 5.3 sits in a profile and does the long grind — drafting, building, fixing, checking, all day.
A judge agent scores every output and sends weak work back around the loop until it actually passes.
What you're watching: a real countdown timer GLM 5.3 built unattended — six steps, three files, checked its own work, then reported back. This is it running.
This is the exact system running my sites — think Trello, except the cards do the work themselves.
What you're watching: the real board inside my Agent OS — switching the team onto the Hermes cluster, typing one topic in, and the Backlog / Building / Review / Done columns it moves through.
A team of GLM 5.3 agents picks the card up, and nothing ships until the judge says it passes.
What you're watching: one topic typed in, and the planner agent — running on GLM 5.3 — breaking it into five real article cards on the board.
Here's a real run on GLM 5.3 — the judge scored the first draft 6 out of 10 and named exactly what was wrong.
What you're watching: the actual writer-and-judge run through Hermes on GLM 5.3. First pass 6/10 — "generic placeholders, no named system". After the rewrite: 9/10.
Honestly, fair. Raw AI output with no checking is often generic — that's exactly what the judge caught above.
You're not trusting one model's first try. Nothing ships until it clears the bar.
I gave the board one topic and let it run. This is what GLM 5.3 handed back — 1,639 words, front matter, keywords and all.
What you're watching: the real article the GLM 5.3 content team produced on this run, previewed straight off the board.
With 5.2 the judge loop worked, but long runs got expensive and agents lost the thread on big jobs.
What you're watching: every run behind this page, logged against the glm-5-3 profile — the writer, the judge scoring the rewrite, the article, and the timer build. 10 sessions, 64 messages.
Claude does the thinking for five minutes. GLM 5.3 does the doing for five hours.
What you're watching: the real handover — Claude wrote the six-step plan, GLM 5.3 executed every step, checked its own files and reported back. That's the timer you saw earlier.
GLM 5.3 edges past Claude Opus 4.8 on Z.ai's coding benchmark. Claude Fable 5 still leads — the very top model is still ahead.
For an execution layer that runs all day on a flat plan, that's frontier territory.
Wrong: "I'm not technical. I can't wire up a team of AI agents."
Right: It's a board with cards on it. You type a topic, the agents pick it up themselves. If you can use Trello, you can run this.
Wrong: "Running teams of agents must cost a fortune in tokens."
Right: It runs on a flat coding plan, and 5.3 uses fewer tokens per task than 5.2 did. Same serving footprint, same base model — it just got smarter.
Wrong: "AI content is generic, so I don't trust any of it."
Right: That's why the judge agent exists. It blocks generic, it blocks made-up facts, and it sends weak drafts back until they pass.
Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.
Read the 158-page wins doc →The weights aren't even public yet — you're looking at this model in its first week, and it already runs on a plan you can start today.
GLM 5.3 is roughly a quarter the size of Kimi K3 and under a third of Qwen 3.8 Max — and it trades blows with both on agentic coding.
The gap between cheap open models and expensive closed ones keeps shrinking on exactly the tasks that matter for business.
Swap the model in. Everything you built gets smarter today, and nothing breaks.
This is the best moment there's been to build one — the worker models are finally good enough and cheap enough to run all day.
Every new model release makes an existing machine better for free. That's the real advantage of starting early.
This page shows you the setup. The Boardroom hands you the finished machine — the agent profiles, the Kanban board, the judge agent and the auto-publish flow, already connected.
It's live on the coding plan now — the open weights land in about two weeks.
Same 743B base means one model swap, and your whole board upgrades.
~50k per task instead of 120k — long runs actually reach the end.
The judge agent scores every draft and loops it until it passes.
Everything on this page, in one run: the profile wired, the manager, the board, the planner, the judge scoring, and the finished post.
Get your system set up now, because the weights drop in two weeks and this is only getting better.