Run all of this inside the full Agent OS — join the AI Profit Boardroom →

agentos.guide › blog

The Best Hermes Agent Model, And The Five Model Lanes I Route Inside My Agent OS

By Julian Goldie · 7 October 2026 · agentos.guide

A glowing Hermes messenger figure at a control desk with five neon lanes labelled Main, Long Loop, Cheap, Free and Local, each ending in a different AI brain chip

The best Hermes Agent model is Claude Opus 5, run through the Claude subscription you already pay for.

My best free pick is Solar Mini 4, my best local pick is Agents-A1, my best cheap pick is MiniMax M3, and my best pick for long agent loops is Kimi K3.

The model is the brain inside Hermes Agent, so it reads each prompt, chooses the next tool and decides when the job is done.

I don't run one model, because I route five lanes inside my Agent OS and send each job to the brain that fits it.

Here is the ranked list first, then a quick-start, then the exact routing.

The best Hermes Agent model picks, ranked

Here is my full ranking, with the lane each model fills in my Agent OS.

RankModelLane in my Agent OSGoldieBench averageCostRoute into Hermes
1Claude Opus 5Main brain (best overall)8.27Your Claude subscriptionOfficial Claude Subscription DirectSDK plugin
2GPT-5.6 SolSecond brain and reviewer8.16$5 in and $30 out per million tokensOpenRouter profile
3Kimi K3Long agent loops7.89Flat Kimi coding planCoding plan profile
4MiniMax M3Cheap automation (best cheap)7.97$0.30 in and $1.50 out per million tokensIts own Hermes profile
5GLM-5.2 and GLM-5.3Value brain and fallback7.77 for GLM-5.2GLM Coding PlanCoding plan profile
6Grok 4.7Fast fixes7.15 on 20 tasks$2 in and $6 out per million tokensOpenRouter profile
7Solar Mini 4Free lane (best free)Not on the boardFree on Nous Portal for a limited timeNous Portal free list
8Agents-A1Local lane (best local)4.83Free on your own machineOllama profile

The GoldieBench numbers are the live averages on 10 October 2026.

GoldieBench scores one-shot builds, where a model gets one prompt and one attempt.

It doesn't score agent loops, so my long-loop pick comes from hands-on use and not from a number.

These rankings are my own picks, based on the models I've actually wired into Hermes.

What the model does inside Hermes Agent

Hermes Agent is an open-source AI agent made by Nous Research.

It gives a model a body, which means tools, memory, skills, schedules and safety checks.

It doesn't ship with a brain of its own, so you choose the model.

Each time you send a message, Hermes builds a prompt from your text, your memory and your tool list.

The model replies with an answer or a tool call.

Hermes runs the tool, passes the result back, and the loop continues until the job is finished.

A better model makes better choices inside that loop.

A cheaper model lets that loop run for longer.

That trade-off is the whole reason this ranking exists.

Quick-start: set up your first three brains

This quick-start assumes Hermes is already installed.

StepWhat you doWhat you get
1Run hermes model and choose a provider and a default model.Hermes has a working brain.
2Create one profile per model, named after the model.Each brain keeps its own settings and memory.
3Install the Claude subscription plugin in your main profile.Claude Opus 5 runs with no API key.
4Add a free profile that uses a free model on Nous Portal.Small jobs cost nothing.
5Add a local profile that points at Ollama.Private jobs stay on your machine.
6Set a fallback provider with hermes fallback.The agent keeps going when the main model fails.
7Open the Hermes tab in your Agent OS and pick a profile.You switch brains from one dropdown.

The first three steps give you the best Hermes Agent model as your main brain.

The next three steps make sure you never pay premium prices for small jobs.

The last step is where the Agent OS earns its keep.

🔥 Want the exact Hermes Agent models in the Agent OS setup I run inside my Agent OS?

Inside the AI Profit Boardroom you get the installable Agent OS, step-by-step video tutorials, four coaching calls a week and 3,400+ members building real automations.

→ Get access here

How the best Hermes Agent model fits an Agent OS

An Agent OS is a dashboard that runs your agents in one place with shared memory.

Mine has a Hermes tab with a profile picker at the top of the chat.

Every Hermes profile shows up in that picker on its own.

When I made a profile for Grok 4.7, I didn't change a single line of the Agent OS.

I made the profile, and it appeared in the list.

That's why I think in lanes and not in single models.

LaneModel I useJobs I send down itWhy
MainClaude Opus 5Planning, client-facing writing and hard buildsIt's the top single model on GoldieBench, and my subscription covers it.
Long loopKimi K3Research, big codebases and jobs with many stepsIt holds about one million tokens and is tuned for long-horizon work.
CheapMiniMax M3Automations that run all dayIt scores 7.97 for $0.30 per million input tokens.
FreeSolar Mini 4News round-ups, quick questions and simple tool tasksIt's fast and free, and it's good enough for basic agent work.
LocalAgents-A1, Gemma 4 12B or LFM2.5Summaries, file reads and private dataNothing leaves my machine and nothing is metered.

Each lane is one profile.

Each job goes down the lane that fits it.

The expensive brain only touches the work that needs it.

Lane 1: Claude Opus 5 as the main brain

Claude Opus 5 is the top single model on GoldieBench at 8.27.

An official plugin from Nous Research, announced on 22 September 2026, lets Hermes use your Claude Code login.

You don't need an API key or a second account.

Hermes stays in charge, so it still decides which tools run and still asks before anything risky.

The plugin needs Hermes 0.21.4 or newer, and mine refused to load until I updated.

One line in the profile decides which Claude you get.

The opus setting gives you Opus, sonnet gives you Sonnet 5, haiku gives you Haiku 4.5 and fable gives you Fable.

Fable draws usage credits unless you're on the Max plan.

My full walkthrough is in the Claude subscription guide.

Lane 2: Kimi K3 for long agent loops

Kimi K3 is Moonshot AI's flagship, and it scores 7.89 on GoldieBench.

It holds about one million tokens of context, which is roughly a whole codebase in view at once.

It's tuned for long-horizon agent work, which means jobs with many steps over many minutes.

If you pay for the Kimi coding plan, K3 is already on your account.

I pointed one Hermes profile at the coding plan and had it running in about two minutes.

DeepSeek's Flash tier is my alternative for this lane, because it's cheap and was retrained for agent loops.

DeepSeek V4 Flash is currently unranked on GoldieBench, so I can't give you a score for it.

I cover that pairing in Hermes + DeepSeek.

Lane 3: MiniMax M3 for cheap automation

MiniMax M3 scores 7.97 on GoldieBench.

It's the cheapest big-context model on the board, at $0.30 per million input tokens and $1.50 per million output tokens.

I use it for automations where the agent calls a lot of tools.

GLM is my other value pick, and GLM-5.2 scores 7.77.

GLM-5.3 is live on the GLM Coding Plan, and moving to it was a one-line change.

GLM sometimes cuts long outputs short, so I keep a fallback behind it.

Lane 4: Solar Mini 4 for free

Solar Mini 4 is a model from Upstage AI in South Korea.

It has 3 billion active parameters out of 35 billion, and a context window of 500,000 tokens.

It's free on Nous Portal for a limited time.

In my test it replied fast and pulled AI automation news from the last seven days with sources.

It isn't frontier level, and it won't replace Claude.

To set it up, open the model settings, type "solar", choose Nous Portal and pick the variant marked free.

There's a paid variant with a similar name, so check before you click.

Free models get rate limited, so I keep a second free model ready.

Nous Portal also lists other free models, and OpenRouter has a rotating free list of its own.

I list every free lane in the free models guide.

🔥 Want the exact Hermes Agent models in the Agent OS setup I run inside my Agent OS?

Inside the AI Profit Boardroom you get the installable Agent OS, step-by-step video tutorials, four coaching calls a week and 3,400+ members building real automations.

→ Get access here

Lane 5: local models for private work

A local model runs on your own computer through Ollama.

Agents-A1 is my pick for agent loops, because it's tuned for tool calling and runs at about 95 tokens a second on my 36GB Mac.

It scores 4.83 on GoldieBench, so it isn't a strong one-shot builder.

Qwable 5 27B Coder scores 7.14, which is the best build quality of any local model, but it's slow for loops.

Gemma 4 12B on the MLX path scores 3.98 and runs comfortably on 16GB.

LFM2.5 from LiquidAI fits on an 8GB laptop and ran at more than 140 tokens a second for me.

LFM2.5 isn't on GoldieBench.

In my testing it handled tool calls and a five-step tool chain, but it failed to build full web apps.

That tells you what local models are for.

Give them agent duty, and send the big builds to a bigger brain.

My setup notes are in Run Hermes Free Forever.

Two extra brains I keep on the bench

GPT-5.6 Sol scores 8.16 on GoldieBench, which makes it the closest rival to Claude Opus 5.

I run GPT models in Hermes through an OpenRouter profile.

I use it as a reviewer, because a second flagship catches mistakes the first one misses.

Grok 4.7 scores 7.15 across 20 tasks, so it has a smaller sample behind it.

I wired it in through OpenRouter on the day I tested it.

It built a working kanban app in 82 seconds with four model calls, and that run cost about $0.13.

It also fixed and tested a small bug in 13.9 seconds, while my default profile took 35.6 seconds.

Both got the fix right, so the difference was speed.

You can read that test in the Grok 4.7 Hermes guide.

What each lane costs to run

The main lane costs nothing extra if you already pay for Claude.

The long-loop lane is a flat fee on the Kimi coding plan, so a long job doesn't change the bill.

The cheap lane is pay per token, and MiniMax M3 is priced low enough that I stop watching the meter.

The free lane costs nothing, but it can be rate limited and the free list changes.

The local lane costs nothing per run, but it needs enough memory on your own machine.

I don't state a total, because your bill depends on which plans you already hold.

Tips for running more than one model

Check the real model with hermes -p yourprofile status, because a profile label can be wrong.

Don't trust a model to tell you its own name, because Kimi K3 told me it was an older model.

When you clone a profile, rewrite its personality file, because it copies the old model's notes.

Keep free and local brains as your default, and call the big brain only when the job needs it.

Re-test when a new version lands, because Claude Opus 5.5 scores 7.57 against 8.27 for Opus 5 on one-shot builds.

Common mistakes

The first mistake is running every job on your most expensive model.

The second mistake is picking the paid variant of a free model by accident.

The third mistake is having no fallback, so one outage stops the whole agent.

The fourth mistake is reading a one-shot build score as proof of agent-loop quality.

The fifth mistake is giving a small local model a job that needs a frontier brain.

Where to go next

Once your models are set, the rest of the stack matters just as much.

My best Hermes Agent setup post covers the full install and layout.

My best Hermes Agent memory post covers how any model remembers your work.

I've also written an earlier list of Hermes agent models on the AI Profit Boardroom blog, which covers eight model families.

Also On Our Network

FAQ

What is the best Hermes Agent model?

Claude Opus 5 is the best Hermes Agent model for most people. It is the top single model on GoldieBench at 8.27, and an official Nous Research plugin runs it inside Hermes on a Claude subscription with no API key.

How do I switch models in Hermes Agent?

Run hermes model to choose a provider and a default model. I keep one profile per model and switch with the -p flag, and every profile appears in the Agent OS chat picker automatically.

What is the best free model for Hermes Agent?

Solar Mini 4 from Upstage AI is my best free pick while it is free on Nous Portal. It is fast and handles basic agent tasks, but it is not frontier level and it is not on GoldieBench.

What is the best local model for Hermes Agent?

Agents-A1 is my best local pick for agent loops, because it is tuned for tool calling and runs at about 95 tokens a second on a 36GB Mac. Gemma 4 12B suits 16GB machines and LFM2.5 fits on 8GB.

Does GoldieBench test agent loops?

No. GoldieBench scores one-shot builds, where each model gets one prompt and one attempt. It is a good guide to build quality, but my long-loop pick comes from hands-on use.

Final word on the best Hermes Agent model

Put Claude Opus 5 in your main lane, add a free lane and a local lane beside it, and set a fallback.

Route each job to the brain that fits, and you'll get the most from the best Hermes Agent model.

About Julian

I'm Julian Goldie, an AI entrepreneur, SEO expert and the founder of the AI Profit Boardroom, which has 3,400+ members.

I help business owners scale with AI agents, automation and SEO.

My YouTube channel has 400,000+ subscribers, and I run Goldie Agency, a seven-figure SEO agency.

→ Get my best AI training inside the AI Profit Boardroom

Related guides on agentos.guide

→ Run Claude Opus 5 Inside Hermes On Your Claude Subscription

→ The Kimi K3 Machine

→ Every Free Model Lane For Hermes

→ Run Hermes Free Forever

Get the whole Agent OS

The installable system behind every guide on this site — dashboard, agents, memory and pipelines — updated as new tools land, with 3,400+ members and four coaching calls a week.

Join the AI Profit Boardroom →

📺 Video notes + links to the tools 👉

🎥 Learn how I make these videos 👉

🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉

← all guides on agentos.guide