GLM 5.3 + Hermes Agent · 14 August 2026

Same brain. Way better training.

GLM 5.3 plus Hermes Agent is one of the most powerful AI agent setups you can run right now.

Z.ai changed nothing about the model's brain — not one new parameter.

And it still jumped six times higher on terminal tasks, and beat Claude Mythos 5 and GPT-5.6 Sol on a benchmark that shocked people.

It got so good at one thing that Z.ai is holding the open weights back.

There's one number buried in this release that matters more than any score — and almost nobody is talking about it.

743Bsame base model
6.2×terminal tasks
84.5CyberGym — 1st
~50ktokens per task
§1 Straight from Z.ai

Read it, and run it, yourself.

Every number on this page comes from Z.ai's own release — nothing second-hand.

Official Z.ai sources + resources ↓

"GLM-5.3 runs on the same mixture-of-experts base as GLM-5.2, and every reported gain came from extended post-training rather than a fresh pretrain."

— Z.ai's GLM-5.3 release, 14 August 2026

§2 What actually changed

The brain didn't get bigger. The brain went to school.

Same 743 billion parameters as GLM 5.2. Every single gain came from a month of extra training.

GLM 5.2 743B the base model post-training 1 month inside fake work environments GLM 5.3 743B identical. zero new. Not one new parameter — the whole jump is training.

The base is untouched. Everything that changed sits in the middle box.

§3 Where it went to school

They trained it on jobs that take days.

Not quick questions. Long, multi-step work — the kind an experienced engineer needs days to finish.

read plan build check fix done Some training environments represent days of work for an experienced engineer.

This is the shape of the work it practised on — long chains, not single answers.

§4 The problem

The Halfway Problem.

That's exactly the kind of work you want an agent to do.

Not answer one question. Finish a whole project.

WITHOUT LONG-JOB TRAINING drifts off · loses the plot · gives up starthalfwaydone TRAINED ON LONG, MULTI-STEP JOBS Finishing is the whole skill.
Thinking it? "My agent always dies before the job is done."

That's the exact failure this release went after.

The model was trained inside long jobs, so holding the thread is the thing it practised most.

The brain didn't get bigger. It just learned how to finish.
§5 The numbers · terminal work

Terminal-Bench 3.0: 4.6 → 28.3

This tests how well a model works inside a computer terminal on its own. Roughly a six times jump — same brain.

4.6 GLM 5.2 28.3 GLM 5.3 6.2×
§6 The numbers · coding

About fifty percent better at coding.

On Z.ai's own coding benchmark, GLM 5.3 hits 34.5% at its highest effort setting. GLM 5.2 got 23.4%.

GLM 5.2 23.4% GLM 5.3 34.5% Z.ai Code Bench, highest effort setting — Z.ai's own reported figures.
§7 The numbers · agent work

Best scores on the boards built for agents.

Agents' Last Exam and DeepSWE both test whether a model can actually run a job, not just answer about one.

23.8 28.5 Agents' Last Exam (CLI) 46.2 66.9 DeepSWE v1.1 GLM 5.2 GLM 5.3
§8 The one that made headlines

It went ahead of Claude Mythos 5 and GPT-5.6 Sol.

CyberGym is a UC Berkeley benchmark: can a model find and confirm real bugs in real source code?

GLM 5.3 84.5 Claude Mythos 5 83.8 GPT-5.6 Sol 83.6 An open-weights model went ahead of the two biggest closed models on this test.
§9 During training

2,436 real bugs. 269 real projects.

More than a thousand of them were rated critical or high.

2,436 confirmed bugs found 269 open source projects 1,097 critical or high
§10 Quick honesty check

These are Z.ai's own numbers.

They haven't been independently verified yet — and on most other security tests, GLM 5.3 still sits well behind.

EXPLOITBENCH — THE OTHER DIRECTION GLM 5.3 54.4% Mythos 5 78.0% Finding flaws is not the same as exploiting them. Here Claude is comfortably ahead.

Worth saying out loud so nothing here gets overstated.

§11 The decision they've never made before

They're holding the weights back.

GLM 5.2 shipped its weights within days. This one didn't — roughly two weeks of safety checks first.

14 Aug model live on the coding plan ~2 weeks of safety checks ~28 Aug open weights expected

When a lab famous for shipping everything fast decides to slow down, that tells you something.

§12 But here's the thing

You don't have to wait.

GLM 5.3 is live right now on the GLM Coding Plan. And that's exactly where the Hermes setup comes in.

Old waywait 2 weeks
  • Wait for the open weights to drop
  • Find a machine big enough to serve 743B
  • Work out serving, quantisation, throughput
  • Rebuild your agent wiring around it
  • Start learning it from zero, two weeks late
New way~5 min
  • It's already live on the flat coding plan
  • Clone your existing Hermes profile
  • Change one line: the model name
  • Every agent on your board runs on it
  • Swap the weights in later, for free

What you're watching: the whole migration — one line in the config goes from glm-5.2 to glm-5.3. Skills, keys, MCP servers and memory all carry over untouched.

Skip the setup

Get the whole agent system built for you.

You can wire this up yourself with the steps below. Or get the finished thing — the Agent OS, with GLM 5.3, Hermes and your other models already connected.

The Agent OS zip — GLM 5.3 ready to wire in through Hermes
The Kanban board + judge agent, already built, already looping
Every CLI you already pay for in one dashboard — Claude, Codex, OpenClaw, Free Claude Code
A 30-day roadmap and video tutorial for setting the whole thing up
4 live coaching calls a week — bring your own GLM setup and get unstuck
Daily tutorials — including wiring GLM 5.3 into Hermes step by step
4,000+ business owners across 38 countries, plenty already running GLM models
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Set up in an afternoon · used in 38 countries · new tools added the week they ship
§13 How it works inside Hermes

GLM 5.3 becomes its own agent profile.

Hermes lets you plug in any model and give it a profile. Here's the real setup on my machine — clone, change one line, run.

What you're watching: the real session — cloning the GLM 5.2 profile, pointing it at glm-5.3, then running a live job that writes a file and runs a shell command.

§14 Sitting next to your other agents

Pick it in the model manager. That's it.

Because 5.3 shares the exact same base, everything you already had wired up just carries over.

What you're watching: the real Hermes manager with glm-5-3 selected — dropping open to show it sitting beside every other agent profile, with its own skills, keys and config.

Thinking it? "So I have to rebuild everything for the new model?"

No. Same base model means same wiring.

You swap the model name and instantly get the smarter version. Nothing breaks.

Swap the model. Keep the whole machine.
§15 The number nobody's talking about

It does more work with fewer words.

GLM 5.3 scores higher than Claude Opus 4.8 on Z.ai's coding benchmark while using well under half the output tokens.

SCORE → GLM 5.3 31.4% Opus 4.8 29.5% TOKENS BURNED → GLM 5.3 ~50,000 Opus 4.8 120,000 2.4× less

Higher score, top row. Fewer tokens to get there, bottom row. That gap is the whole story.

§16 Why that matters so much

Agents don't answer once. They loop.

They write, check, fix, write again — for hours. Every wasted token is wasted time and wasted plan limits.

write check fix repeat for hours TOKEN-HUNGRY stalls — out of room LEAN More done per token means your long runs actually finish.
Thinking it? "Running teams of agents must cost a fortune in tokens."

It's the opposite, and 5.3 made it more true — this runs on a flat coding plan, and the model now uses fewer tokens per task than the old version did.

Inside the Agent OS the everyday work also runs on free local models and the CLIs you already pay for, so you're not paying twice. There are full token-efficiency tutorials in the Boardroom too.

§17 Why the pairing works

Hermes gives the structure. GLM 5.3 does the work.

A worker that's cheap to run, doesn't waste tokens, and was literally trained on long multi-day tasks.

i.The structure

Hermes holds the profiles, the skills, the memory and the board. It decides who does what, and when.

ii.The worker

GLM 5.3 sits in a profile and does the long grind — drafting, building, fixing, checking, all day.

iii.The referee

A judge agent scores every output and sends weak work back around the loop until it actually passes.

What you're watching: a real countdown timer GLM 5.3 built unattended — six steps, three files, checked its own work, then reported back. This is it running.

§18 What it looks like in practice

Drop one topic on a board. Walk away.

This is the exact system running my sites — think Trello, except the cards do the work themselves.

What you're watching: the real board inside my Agent OS — switching the team onto the Hermes cluster, typing one topic in, and the Backlog / Building / Review / Done columns it moves through.

§19 The team on the card

Researcher. Writer. Editor. Judge.

A team of GLM 5.3 agents picks the card up, and nothing ships until the judge says it passes.

the cardone topic researcherfinds the angle writerdrafts the post editortightens it up judgescores itout of 10 weak or unchecked → sent straight back only then: publish

What you're watching: one topic typed in, and the planner agent — running on GLM 5.3 — breaking it into five real article cards on the board.

§20 The judge loop, for real

It loops until the work actually passes.

Here's a real run on GLM 5.3 — the judge scored the first draft 6 out of 10 and named exactly what was wrong.

What you're watching: the actual writer-and-judge run through Hermes on GLM 5.3. First pass 6/10 — "generic placeholders, no named system". After the rewrite: 9/10.

Thinking it? "AI content is generic. I don't trust it."

Honestly, fair. Raw AI output with no checking is often generic — that's exactly what the judge caught above.

You're not trusting one model's first try. Nothing ships until it clears the bar.

§21 What comes out the end

A finished post, while you're at the gym.

I gave the board one topic and let it run. This is what GLM 5.3 handed back — 1,639 words, front matter, keywords and all.

What you're watching: the real article the GLM 5.3 content team produced on this run, previewed straight off the board.

You go to the gym. You come back. The work is done — and checked.
§22 What changed on this loop

Both of the old problems shrink.

With 5.2 the judge loop worked, but long runs got expensive and agents lost the thread on big jobs.

On GLM 5.2loses the thread
  • Long runs got expensive in tokens
  • Agents drifted on big multi-step jobs
  • Terminal work: 4.6 on Terminal-Bench 3.0
  • 23.4% on Z.ai's own coding benchmark
  • More words to reach a lower score
On GLM 5.3holds it
  • Trained specifically on long, multi-step jobs
  • Holds the thread — that's what it practised
  • Terminal work: 28.3 on the same test
  • 34.5% at highest effort on the same benchmark
  • ~50k tokens per task instead of 120k

What you're watching: every run behind this page, logged against the glm-5-3 profile — the writer, the judge scoring the rewrite, the article, and the timer build. 10 sessions, 64 messages.

§23 The smart way to run it

Expensive model plans. Cheap model works.

Claude does the thinking for five minutes. GLM 5.3 does the doing for five hours.

the thinking 5 minutes plan · strategy · outline hand over the plan the doing 5 hours on a flat coding plan The execution layer runs all day. That's where the hours actually go.

What you're watching: the real handover — Claude wrote the six-step plan, GLM 5.3 executed every step, checked its own files and reported back. That's the timer you saw earlier.

§24 Where it sits now

That gap used to be huge. It isn't anymore.

GLM 5.3 edges past Claude Opus 4.8 on Z.ai's coding benchmark. Claude Fable 5 still leads — the very top model is still ahead.

39.5 Claude Fable 5 max effort 34.5 GLM 5.3 highest effort 31.4 GLM 5.3 ~50k tokens 29.5 Claude Opus 4.8 120k tokens

For an execution layer that runs all day on a flat plan, that's frontier territory.

§25 Three things people say every time

Three beliefs to drop.

"I'm not technical" "tokens cost a fortune" "AI content is generic" it's a board of cards it's a flat plan the judge blocks it

Wrong: "I'm not technical. I can't wire up a team of AI agents."

Right: It's a board with cards on it. You type a topic, the agents pick it up themselves. If you can use Trello, you can run this.

Wrong: "Running teams of agents must cost a fortune in tokens."

Right: It runs on a flat coding plan, and 5.3 uses fewer tokens per task than 5.2 did. Same serving footprint, same base model — it just got smarter.

Wrong: "AI content is generic, so I don't trust any of it."

Right: That's why the judge agent exists. It blocks generic, it blocks made-up facts, and it sends weak drafts back until they pass.

Don't take my word for it

Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.

Read the 158-page wins doc →
§26 One more thing people think

"I'll just wait until this settles down."

The weights aren't even public yet — you're looking at this model in its first week, and it already runs on a plan you can start today.

your board built once, stays built GLM 5.2 GLM 5.3 whatever's next My GLM 5.2 board became a GLM 5.3 board with one model swap.
§27 Zoom out

Smaller, cheaper, and competitive.

GLM 5.3 is roughly a quarter the size of Kimi K3 and under a third of Qwen 3.8 Max — and it trades blows with both on agentic coding.

GLM 5.3 743B Qwen 3.8 Max 3× bigger+ Kimi K3 4× bigger

The gap between cheap open models and expensive closed ones keeps shrinking on exactly the tasks that matter for business.

§28 For you, the takeaway

The execution layer just got a serious upgrade.

i.Already have an agent system?

Swap the model in. Everything you built gets smarter today, and nothing breaks.

ii.Don't have one yet?

This is the best moment there's been to build one — the worker models are finally good enough and cheap enough to run all day.

iii.Either way, build the system.

Every new model release makes an existing machine better for free. That's the real advantage of starting early.

today system already built → every release compounds no system → every release starts from zero again
Your move

Get the Agent OS, with GLM 5.3 ready to wire in.

This page shows you the setup. The Boardroom hands you the finished machine — the agent profiles, the Kanban board, the judge agent and the auto-publish flow, already connected.

The Agent OS zip + video tutorial — GLM 5.3 through Hermes, ready to go
Slots for your Claude, OpenClaw, Codex and Free Claude Code, all in one dashboard
Tools I've added for videos, SEO agents and AI avatars
The 30-day roadmap and daily updates as the Agent OS improves with every model drop
When the GLM 5.3 weights go public in the next couple of weeks, you get the updated build
4 weekly coaching calls — bring your GLM 5.3 setup and get help live
A prompt library for every workflow, plus a member map to find people near you
4,000+ business owners, many already running GLM agents for content, leads and client work
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Someone online 24/7 · 38 countries · new tools added the week they ship
§29 That's GLM 5.3 plus Hermes

Same brain, way better training, fewer tokens.

You stopped waiting for weights.

It's live on the coding plan now — the open weights land in about two weeks.

You stopped rebuilding.

Same 743B base means one model swap, and your whole board upgrades.

You stopped burning tokens.

~50k per task instead of 120k — long runs actually reach the end.

You stopped shipping generic work.

The judge agent scores every draft and loops it until it passes.

Everything on this page, in one run: the profile wired, the manager, the board, the planner, the judge scoring, and the finished post.

Get your system set up now, because the weights drop in two weeks and this is only getting better.