The Goldie framework · Hermes agents × Kimi K3

The Long-Haul Engine.

Hermes is the agent runtime in my Agent OS — hands, tools, memory. Kimi K3 is Moonshot's new freight-engine brain — a million tokens, tuned for long unsupervised jobs. One command marries them: kimi-k3 -z "the job" --yolo — the profile's own wrapper command. Below: the real agentic workflows it ran on day one, including making a 3D video by itself.

A colossal ornate brass locomotive engine pulling a long train of glowing carriages across a mountain viaduct at night, driven by a fully clothed keeper in a full-length hooded robe while fully clothed winged lantern-bearers in knee-length tunics light the track ahead
THE LONG-HAUL ENGINE — runtime + brain + tools = artifacts Hermes hands, tools, memory Kimi K3 1M-token brain MCP tools Blender · files · web Artifacts video · docs · sitesone wrapper command dispatches it: kimi-k3 -z "the job" --yolo · then you walk away
the marriage in one line — Hermes supplies the hands, K3 supplies the attention span
1M
tokens of working memory per run
1
profile to hire it as an agent
100%
of the demo task self-verified
$0
extra — rides the Kimi coding plan
✦
I · the problem

The Sprinter Problem.

You've delegated to AI agents before. You know how it ends.

Step one: brilliant. Step two: good.

Step five: it forgot what step one was for.

Step eight: it declares victory on a job it half did.

Most models are sprinters — dazzling for thirty seconds, useless over distance.

So you stay in the loop. You check every step. You are the project manager of a very fast intern.

Delegation that needs constant supervision isn't delegation.

The Long-Haul Engine breaks that cycle for good.

Thinking it?"Every model claims to be 'agentic' now. Why believe this one?"

Two concrete reasons: the launch benchmarks' standout is Terminal Bench — the test of driving tools over many steps — and the demo run below ends with the agent VERIFYING its own output before reporting. Claims are cheap; verified file writes aren't.

✦
II · how it works, in simple words

A freight engine, hired as staff.

Hermes is Nous Research's agent runtime — the backbone of my Agent OS. It gives whatever model you choose real HANDS: the terminal, file access, web tools, and the whole MCP catalog (hermes mcp install blender gave my agents 22 Blender tools last week). Models plug in as profiles; my stack runs 36 of them.

K3 is Moonshot's new flagship brain: 2.8 trillion parameters (Moonshot's docs corrected the launch-day 2.8T estimate), a one-million-token context, tuned for long-horizon work. On the Kimi coding plan it costs nothing extra — hermes profile create kimi-k3 --clone-from kimi-highspeed and it's staff.

The marriage matters because each half fixes the other's gap. Hermes without a long brain forgets mid-job; K3 without hands can only talk. Together: an agent that holds the WHOLE job in memory while actually operating your tools — Blender, the terminal, your files, your workflows.

The rule of thumb this guide keeps returning to: give K3 the jobs where being RIGHT at the end matters more than being fast in the middle.

A sprinter model over a 10-step job drifts… K3 over the same job done ✓ the long-haul trade: slower per step, flat over distance — the finish line is the point
why "slow" is the wrong complaint — coherence over steps is the metric that pays
Thinking it?"A million tokens of context — does that actually matter for agents?"

It's the whole game. Agents die when the job outgrows their memory — step twelve can't see step one's requirements anymore.

I tested the window directly on launch day: a fact buried in 162,000 tokens of noise, recalled exactly in 18 seconds. For an agent, that means the brief NEVER falls out of view.

✦
III · exactly how it works — real agentic workflows

Three real Hermes + K3 workflows, witnessed.

Not hypotheticals — each of these ran on my machine through the kimi-k3 Hermes profile wrapper, and each artifact is embedded or linked in this guide.

Workflow 1 · The agent makes a video (Hermes + K3 + Blender MCP)

The showcase: one prompt asked the K3 agent to build a rocket-launch scene in Blender, animate a 150-frame liftoff, and render it to MP4 — modelling, keyframing, camera work and encoding, all driven tool-call by tool-call through Hermes's Blender MCP connection.

The agent's render · 150 frames, modelled + animated + encoded by the K3 agent through Blender

What you're looking at: the actual MP4 the agent saved to disk — scene modelling, lighting, 150 keyframed frames and H264 encoding, every step a tool call through Hermes's Blender MCP connection. No human touched Blender.

Workflow 2 · The verified deliverable (write → check → report)

A complete job with a checkable end state — not a question, a deliverable. Here's that run:

1

The hire. hermes profile create kimi-k3 --clone-from kimi-highspeed, model pointed at k3 on the Kimi coding endpoint, then hermes profile alias kimi-k3 for the wrapper command. One profile — K3 now takes Hermes jobs like any other staff member.

2

The brief. "Create a file at this exact path containing a 300-word briefing with 5 specific numbered steps. Write it, verify it exists, report the word count." Note the shape: destination + cargo + proof.

3

The run. One command: kimi-k3 -z "…" --yolo. No babysitting, no follow-ups. K3 planned, wrote and structured the whole document itself.

4

The self-check. It didn't just claim success — it checked the file on disk and reported back: "File exists on disk (2,119 bytes, permissions confirmed)." The agent audited itself before telling me it was done.

5

The cargo. The briefing itself is genuinely sharp — embedded complete below, verbatim:

k3-agent-demo.md · written + verified by the K3 agent · 2,119 bytes ● unedited, complete
# Launch-Day Playbook: adding a brand-new AI model to a production stack in one afternoon

A new model card drops at 9 a.m. and the team wants it live by end of day. That timeline is fine — if you refuse to improvise. This playbook assumes a stack that already serves several models behind profiles, a benchmark harness, and real jobs waiting in a queue.

1. Check what your existing plans already include. Before touching config, read your current rollout and incident plans. A mature stack already has a new-model checklist covering quotas, fallback behavior, cost ceilings, and rollback triggers. Reusing it is not laziness; it is how you avoid re-learning last quarter's outage. Note what the plan misses — new context lengths, new tool-call formats, different rate limits — and patch only those gaps.

2. Verify the served model via API, not self-report. Never trust an announcement page, a chat reply claiming an identity, or a colleague's screenshot. Hit the endpoint directly: list the model registry, send a fixed probe prompt, and confirm response metadata and tokenizer behavior match the model you intend to run. Self-reports lie; endpoints do not.

3. Wire it as one profile beside existing models. Do not rip anything out. Add the new model as an additional profile with its own credentials, limits, and routing rules, leaving the incumbent untouched as default. Side-by-side deployment makes every comparison cheap and rollback a config toggle instead of a redeploy.

4. Run one real benchmark before trusting it. Pick the benchmark your team already respects — the one correlated with your actual workload, not a public leaderboard. One honest run on real prompts beats a day of vibes. Record the numbers beside the incumbent's so the decision stays visible to everyone.

5. Give it the long-horizon jobs first. Counterintuitive but correct: long agentic runs expose context drift, tool-call flakiness, and cost blowouts that short chats hide. Assign a bounded batch of your longest real jobs, watch them end to end, then widen.

Total: one afternoon spent, one model launched, sleep preserved.

The actual file K3's agent wrote and verified — unedited opening. Note the tone: it writes like an operator, not a chatbot.

Workflow 3 · K3 as staff inside bigger Hermes machinery

Because K3 is a named Hermes profile, every Hermes workflow can hire it without new wiring: the Agent Kanban's Planner→Builder→Reviewer teams can run it as the Builder, the Loop Engine can use it as builder or judge, and any cron or webhook job can call --profile kimi-k3. One profile, every workflow in the OS — that's the Hermes model: wire once, hire everywhere.

So when someone asks "but what IS it?" — it's a model that treats a job like a train route: load everything at the start, stop at every station, confirm arrival — plugged into a runtime that owns the track.

IV · open everything

The engine's paperwork, all real.

Everything from this guide's runs ↓
The Agent OS Kimi tab showing K3 wired in as a selectable mode next to the other Kimi models

What you're looking at: K3 living inside the Agent OS — the same flagship that runs these agent jobs is one switch among the rest of the staff.

Delegation that needs supervision isn't delegation. It's typing with extra steps.
Typical model · 128k tokens Kimi K2.6 · 256k tokens Kimi K3 · 1,000,000 tokens · 2.8T params
the manifest: how much cargo the engine loads before it departs
V · the framework

The five cars of the Long-Haul Engine.

i.

The Full Manifest

Load EVERYTHING at departure — the whole repo, all the docs, the complete brief. A million tokens means nothing gets left on the platform.

ii.

The Destination Brief

Every job ends in a checkable state: a file at a path, a count, a passing test. Destination + cargo + proof — never "help me with…".

iii.

The Slow Boiler

Accept the trade: minutes of thinking per step, coherence across all of them. Schedule it like freight — overnight and unattended.

iv.

The Arrival Check

The agent verifies its own delivery before reporting — file exists, bytes counted. Bake the check INTO the brief.

v.

The Timetable

K3 doesn't replace your fast models — it takes the long routes. Quick chat stays on the sprinters; the freight goes on the engine.

✦
VI · old way vs new way

Delegating to a sprinter vs loading the engine.

Old way
you babysit every step
  • Break the job into ten prompts because the model can't hold it
  • Re-paste context at every step as the window overflows
  • Catch the step-7 drift where it forgot the step-1 requirement
  • Trust "Done!" and discover the file was never written
  • Stay at the desk for the whole run
  • Conclude AI agents "aren't there yet"
New way
one brief, verified delivery
  • One brief with the WHOLE job — the million-token manifest takes it
  • Destination + cargo + proof written into the prompt
  • The engine runs unattended; slowness is the schedule, not a problem
  • Self-verification before it reports — bytes on disk, counted
  • You review a finished deliverable, not a stream of steps
  • Fast models keep the quick jobs; the engine takes the long ones
Thinking it?"What long jobs do I even have?"

The ones you've stopped delegating because agents kept fumbling them: the 40-file refactor, the full-archive research summary, the write-verify-deploy chain, the overnight content batch.

You didn't lack long jobs. You lacked an engine that survives them.

Skip the setup

Get the Long-Haul Engine built for you.

You can wire the profile yourself with the SOP below. Or get the whole railway — the Agent Operating System with Hermes, the profiles, and every agent workflow already laid.

The full Agent OS zip — Hermes + the kimi-k3 profile pattern pre-wired
The agent brief templates — destination + cargo + proof, ready to fill
Coaching calls where we run your first long-haul jobs together
A room of 3,900+ operators sharing what their agents shipped
The prompts, the SOPs, and a member map for your city
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Set up in an afternoon · used in 38 countries · new tools added the week they ship
Thinking it?"Doesn't running the Agent OS burn a fortune in tokens?"

No — that's the biggest myth about it. The everyday 90% runs on free local models on your own machine, and free APIs slot in for more — and this guide's engine rides a coding plan that was already paid for.

For the frontier work, the Agent OS drives the plans and CLIs you already own — Claude's CLI with your Claude subscription, K3 with the Kimi plan. A layer on top of what you own, never a second meter.

And inside the AI Profit Boardroom there are full token-optimisation tutorials, so usage drops even further.

✦
Kimi K2.7 Code · $0.75 / M input Kimi K2.6 · $0.95 / M input Kimi K3 · $3 / M input · included in the coding plan
freight rates — the engine's open-market price vs riding the plan you already pay
VII · three beliefs to drop

What's actually holding you back.

Wrong: "A slow model is a worse model."

Right: For unattended work, per-step speed is nearly irrelevant — the job runs while you sleep either way. What you pay for is arriving WRONG. Slow-and-right is the cheaper trade on every long route.

Wrong: "Agents lie about finishing; you can never trust the report."

Right: Agents that VERIFY don't need trusting. Put the check in the brief — file exists, count reported, test passes — and the report becomes evidence instead of a claim.

Wrong: "Long-horizon agent work is a research demo, not a business tool."

Right: The demo in this guide wrote and verified a real deliverable through a production agent runtime on launch day. The research demo phase ended; the routing decision is yours now.

Don't take my word for it

Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.

Read the 158-page wins doc →
✦
18s 162K-TOKEN RECALL, EXACT 5/5 BRIEF STEPS DELIVERED 2,119 BYTES, SELF-VERIFIED 1 COMMAND, ZERO BABYSITTING
the demo run's numbers — every one checkable in the file above
VIII · the receipts

Who's already running this.

The agent file in this guide is on the site, unedited, byte count and all. The profile that made it sits in the same Agent OS that runs the agency, the channel and every guide here.

3,900+founders inside AIPB
400kYouTube subscribers
38countries · live members
163kX followers
Members write up their wins in a 158-page doc — read it here →
✦
Thinking it?"--yolo mode on a big model sounds dangerous."

Scope it: long-haul jobs run in a project folder the agent owns, writing files you review before anything ships. Autonomy over the sandbox, not over production.

The freight brief's "proof" step is also your audit trail — every delivery is checkable.

IX · the SOP

Your first long-haul run tonight.

1

Hire the engine. Clone your closest profile, point it at k3. Two minutes.

2

Pick a job you'd normally babysit. Multi-step, checkable outcome, no taste calls mid-route.

3

Write the freight brief. Destination (exact path/state) + cargo (what, precisely) + proof (verify and report).

4

Load the full manifest. Paste EVERYTHING relevant. Don't ration — the window fits your whole world.

5

Dispatch and walk away. kimi-k3 -z "…" --yolo. Walking away is the discipline.

6

Audit the arrival. Check the verified deliverable — not the transcript. Judge the cargo, not the journey.

7

Grade the route. Would a sprinter have made it? If yes, give the route back. If no, K3 owns it now.

8

Build the timetable. Queue the long routes for nights and weekends. Freight runs best while you're gone.

✦
X · the 30-day roadmap

From babysitter to dispatcher.

Week 1

First verified run

One long job, one freight brief, one self-checked delivery. Feel the trade.

Week 2

Map your routes

List every job you stopped delegating because agents fumbled it. That's the timetable.

Week 3

Night freight

Queue the heaviest job for overnight. Review the delivery over coffee.

Week 4

Split the fleet

Sprinters on chat and quick edits, the engine on everything long. Route by lane, permanently.

Sprinters win demos. Freight engines win businesses.
the recap

What you just gained.

i.
You gained an agent that finishes.

Long-horizon tuning + a million-token manifest — the brief never falls out of view.

ii.
You gained verified delivery.

Bytes on disk, counted and confirmed, before the agent says "done."

iii.
You gained the freight brief.

Destination + cargo + proof — the prompt shape that makes any agent honest.

iv.
You gained the routing rule.

Slow-and-right for long jobs, fast models for the rest. No more one-model religion.

v.
You gained it in one profile.

Two minutes of wiring inside a stack you already have — or the zip that includes it.

vi.
You gained your evenings.

The engine hauls while you're gone. Dispatchers sleep; babysitters don't.

Your move

Run the Long-Haul Engine — or keep babysitting sprinters.

This guide gave you the wiring, a verified real run, and the freight-brief pattern. The Boardroom gives you the whole railway — the Agent OS with Hermes and every profile pre-laid, live calls, and a room full of dispatchers comparing routes.

Readers nod at the metaphor. Operators load a manifest tonight and wake up to delivered cargo.

Decide which one you are tonight.

Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
3,900+ founders · 38 countries · live coaching calls every week · everything from this guide, pre-wired