Hermes is the agent runtime in my Agent OS — hands, tools, memory. Kimi K3 is Moonshot's new freight-engine brain — a million tokens, tuned for long unsupervised jobs. One command marries them: kimi-k3 -z "the job" --yolo — the profile's own wrapper command. Below: the real agentic workflows it ran on day one, including making a 3D video by itself.

You've delegated to AI agents before. You know how it ends.
Step one: brilliant. Step two: good.
Step five: it forgot what step one was for.
Step eight: it declares victory on a job it half did.
Most models are sprinters — dazzling for thirty seconds, useless over distance.
So you stay in the loop. You check every step. You are the project manager of a very fast intern.
Delegation that needs constant supervision isn't delegation.
The Long-Haul Engine breaks that cycle for good.
Two concrete reasons: the launch benchmarks' standout is Terminal Bench — the test of driving tools over many steps — and the demo run below ends with the agent VERIFYING its own output before reporting. Claims are cheap; verified file writes aren't.
Hermes is Nous Research's agent runtime — the backbone of my Agent OS. It gives whatever model you choose real HANDS: the terminal, file access, web tools, and the whole MCP catalog (hermes mcp install blender gave my agents 22 Blender tools last week). Models plug in as profiles; my stack runs 36 of them.
K3 is Moonshot's new flagship brain: 2.8 trillion parameters (Moonshot's docs corrected the launch-day 2.8T estimate), a one-million-token context, tuned for long-horizon work. On the Kimi coding plan it costs nothing extra — hermes profile create kimi-k3 --clone-from kimi-highspeed and it's staff.
The marriage matters because each half fixes the other's gap. Hermes without a long brain forgets mid-job; K3 without hands can only talk. Together: an agent that holds the WHOLE job in memory while actually operating your tools — Blender, the terminal, your files, your workflows.
The rule of thumb this guide keeps returning to: give K3 the jobs where being RIGHT at the end matters more than being fast in the middle.
It's the whole game. Agents die when the job outgrows their memory — step twelve can't see step one's requirements anymore.
I tested the window directly on launch day: a fact buried in 162,000 tokens of noise, recalled exactly in 18 seconds. For an agent, that means the brief NEVER falls out of view.
Not hypotheticals — each of these ran on my machine through the kimi-k3 Hermes profile wrapper, and each artifact is embedded or linked in this guide.
The showcase: one prompt asked the K3 agent to build a rocket-launch scene in Blender, animate a 150-frame liftoff, and render it to MP4 — modelling, keyframing, camera work and encoding, all driven tool-call by tool-call through Hermes's Blender MCP connection.
What you're looking at: the actual MP4 the agent saved to disk — scene modelling, lighting, 150 keyframed frames and H264 encoding, every step a tool call through Hermes's Blender MCP connection. No human touched Blender.
A complete job with a checkable end state — not a question, a deliverable. Here's that run:
The hire. hermes profile create kimi-k3 --clone-from kimi-highspeed, model pointed at k3 on the Kimi coding endpoint, then hermes profile alias kimi-k3 for the wrapper command. One profile — K3 now takes Hermes jobs like any other staff member.
The brief. "Create a file at this exact path containing a 300-word briefing with 5 specific numbered steps. Write it, verify it exists, report the word count." Note the shape: destination + cargo + proof.
The run. One command: kimi-k3 -z "…" --yolo. No babysitting, no follow-ups. K3 planned, wrote and structured the whole document itself.
The self-check. It didn't just claim success — it checked the file on disk and reported back: "File exists on disk (2,119 bytes, permissions confirmed)." The agent audited itself before telling me it was done.
The cargo. The briefing itself is genuinely sharp — embedded complete below, verbatim:
# Launch-Day Playbook: adding a brand-new AI model to a production stack in one afternoon A new model card drops at 9 a.m. and the team wants it live by end of day. That timeline is fine — if you refuse to improvise. This playbook assumes a stack that already serves several models behind profiles, a benchmark harness, and real jobs waiting in a queue. 1. Check what your existing plans already include. Before touching config, read your current rollout and incident plans. A mature stack already has a new-model checklist covering quotas, fallback behavior, cost ceilings, and rollback triggers. Reusing it is not laziness; it is how you avoid re-learning last quarter's outage. Note what the plan misses — new context lengths, new tool-call formats, different rate limits — and patch only those gaps. 2. Verify the served model via API, not self-report. Never trust an announcement page, a chat reply claiming an identity, or a colleague's screenshot. Hit the endpoint directly: list the model registry, send a fixed probe prompt, and confirm response metadata and tokenizer behavior match the model you intend to run. Self-reports lie; endpoints do not. 3. Wire it as one profile beside existing models. Do not rip anything out. Add the new model as an additional profile with its own credentials, limits, and routing rules, leaving the incumbent untouched as default. Side-by-side deployment makes every comparison cheap and rollback a config toggle instead of a redeploy. 4. Run one real benchmark before trusting it. Pick the benchmark your team already respects — the one correlated with your actual workload, not a public leaderboard. One honest run on real prompts beats a day of vibes. Record the numbers beside the incumbent's so the decision stays visible to everyone. 5. Give it the long-horizon jobs first. Counterintuitive but correct: long agentic runs expose context drift, tool-call flakiness, and cost blowouts that short chats hide. Assign a bounded batch of your longest real jobs, watch them end to end, then widen. Total: one afternoon spent, one model launched, sleep preserved.
The actual file K3's agent wrote and verified — unedited opening. Note the tone: it writes like an operator, not a chatbot.
Because K3 is a named Hermes profile, every Hermes workflow can hire it without new wiring: the Agent Kanban's Planner→Builder→Reviewer teams can run it as the Builder, the Loop Engine can use it as builder or judge, and any cron or webhook job can call --profile kimi-k3. One profile, every workflow in the OS — that's the Hermes model: wire once, hire everywhere.
So when someone asks "but what IS it?" — it's a model that treats a job like a train route: load everything at the start, stop at every station, confirm arrival — plugged into a runtime that owns the track.

What you're looking at: K3 living inside the Agent OS — the same flagship that runs these agent jobs is one switch among the rest of the staff.
Load EVERYTHING at departure — the whole repo, all the docs, the complete brief. A million tokens means nothing gets left on the platform.
Every job ends in a checkable state: a file at a path, a count, a passing test. Destination + cargo + proof — never "help me with…".
Accept the trade: minutes of thinking per step, coherence across all of them. Schedule it like freight — overnight and unattended.
The agent verifies its own delivery before reporting — file exists, bytes counted. Bake the check INTO the brief.
K3 doesn't replace your fast models — it takes the long routes. Quick chat stays on the sprinters; the freight goes on the engine.
The ones you've stopped delegating because agents kept fumbling them: the 40-file refactor, the full-archive research summary, the write-verify-deploy chain, the overnight content batch.
You didn't lack long jobs. You lacked an engine that survives them.
No — that's the biggest myth about it. The everyday 90% runs on free local models on your own machine, and free APIs slot in for more — and this guide's engine rides a coding plan that was already paid for.
For the frontier work, the Agent OS drives the plans and CLIs you already own — Claude's CLI with your Claude subscription, K3 with the Kimi plan. A layer on top of what you own, never a second meter.
And inside the AI Profit Boardroom there are full token-optimisation tutorials, so usage drops even further.
Wrong: "A slow model is a worse model."
Right: For unattended work, per-step speed is nearly irrelevant — the job runs while you sleep either way. What you pay for is arriving WRONG. Slow-and-right is the cheaper trade on every long route.
Wrong: "Agents lie about finishing; you can never trust the report."
Right: Agents that VERIFY don't need trusting. Put the check in the brief — file exists, count reported, test passes — and the report becomes evidence instead of a claim.
Wrong: "Long-horizon agent work is a research demo, not a business tool."
Right: The demo in this guide wrote and verified a real deliverable through a production agent runtime on launch day. The research demo phase ended; the routing decision is yours now.
Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.
Read the 158-page wins doc →The agent file in this guide is on the site, unedited, byte count and all. The profile that made it sits in the same Agent OS that runs the agency, the channel and every guide here.
Scope it: long-haul jobs run in a project folder the agent owns, writing files you review before anything ships. Autonomy over the sandbox, not over production.
The freight brief's "proof" step is also your audit trail — every delivery is checkable.
Hire the engine. Clone your closest profile, point it at k3. Two minutes.
Pick a job you'd normally babysit. Multi-step, checkable outcome, no taste calls mid-route.
Write the freight brief. Destination (exact path/state) + cargo (what, precisely) + proof (verify and report).
Load the full manifest. Paste EVERYTHING relevant. Don't ration — the window fits your whole world.
Dispatch and walk away. kimi-k3 -z "…" --yolo. Walking away is the discipline.
Audit the arrival. Check the verified deliverable — not the transcript. Judge the cargo, not the journey.
Grade the route. Would a sprinter have made it? If yes, give the route back. If no, K3 owns it now.
Build the timetable. Queue the long routes for nights and weekends. Freight runs best while you're gone.
One long job, one freight brief, one self-checked delivery. Feel the trade.
List every job you stopped delegating because agents fumbled it. That's the timetable.
Queue the heaviest job for overnight. Review the delivery over coffee.
Sprinters on chat and quick edits, the engine on everything long. Route by lane, permanently.
Long-horizon tuning + a million-token manifest — the brief never falls out of view.
Bytes on disk, counted and confirmed, before the agent says "done."
Destination + cargo + proof — the prompt shape that makes any agent honest.
Slow-and-right for long jobs, fast models for the rest. No more one-model religion.
Two minutes of wiring inside a stack you already have — or the zip that includes it.
The engine hauls while you're gone. Dispatchers sleep; babysitters don't.