Pi Agent plus DeepSeek just beat Claude Code. And Hermes. And almost every big-name agent tool.
Same AI model in every single test. Same 30 tasks.
One tool passed 20 of them. Another passed 15.
Nothing about the AI changed between those runs. Only the tool wrapped around it changed.
By the end, you'll know which setup won, why it won, and the one mistake almost everyone makes when they pick an AI agent.
Same engine. Different car. Completely different drive — that's the whole benchmark.
DeepSeek V4 Flash, 30 hard agentic tasks, eight harnesses. The agent has to go do things, use tools, and finish the job on its own.
Composio builds tooling for AI agents. This thread is round two of their test: same DeepSeek V4 Flash model pushed through Hermes Agent, Pi Agent, Prime Agent and Deep Agents — after Claude Code, Codex, OpenCode and OMP in round one. Pi was the cheapest harness and passed the most tasks.
Claude Code is a car. Hermes is a car. Pi is a car. Same engine, different car, completely different drive.
These are hard agentic tasks — go do things, use tools, finish alone. Two out of every three, on a flash model built for speed, is the best score on the board. The point isn't perfection. The point is the same model scored 15 in a different wrapper.
Fair note: in the overall ranking across all eight, Claude Code and OpenCode were quicker than Pi. On pass rate and efficiency together, Pi came out on top.
Across all eight harnesses, the exact same model scored anywhere from 47% to 67% task success. The AI never got smarter or dumber. Only the wrapper changed.
The tool you wrap around your AI multiplies the result — up or down. Right harness: more reliable AND faster. Wrong harness: more failures, slower runs, same intelligence.
Every extra layer is another place for the agent to get lost. A clean harness gives the model a short path from "here's the task" to "task done."
Never judge a model alone. Test the model-harness pair on your own tasks — because the pairing is the product.
"The same model delivered 47–67% task success, cost $0.019–$0.104 per task, and took 122.7–272.4s median time per task depending on the harness. Benchmark the model-harness pair you'll actually use, not the model in isolation."
— Composio (@composio), benchmark thread, Aug 10, 2026
The Agent OS plugs your agents — Claude, Hermes, OpenClaw, Free Claude Code — into one system with shared memory. Run the same task through different setups and see which pairing actually works for your business.
That's the biggest myth about it. The everyday 90% runs on free local models on your own machine, free APIs slot in for more, and for frontier work it drives the CLIs you already pay for — your Claude subscription already includes the Claude Code CLI, and the Agent OS plugs straight into it, so you're never paying twice.
And this benchmark just proved the cheap route can win: the best harness here cost 2.8 cents per finished task.
A fresh, vanilla install. The only thing bolted on was the MCP server plugin — the standard connector that lets an agent plug into outside tools. No tuning. No special configuration.
What you're watching: the real install from this morning, replayed at reading pace — one npm command, 3 seconds, and Pi 0.84.1 is live with DeepSeek V4 Flash available. The exact winning pair, on my machine.
Then I gave that vanilla install one prompt. Real session, replayed at reading pace: Pi + DeepSeek V4 Flash wrote a 504-line neon snake game in 56 seconds.
And here's that exact game running — real screen recording, played by an autopilot script. Score ticks up, particles fire, game over screen works. One prompt, zero fixes.
Everything above ran on the out-of-the-box install you just watched. That's the point of this whole benchmark: the vanilla setup beat the heavily-engineered ones.
Sessions up to 3.5 million tokens and 33 tool calls — like an agent writing itself a to-do list the length of a phone book before starting the job.
Wrong: "I need the most powerful, top-tier model to get good results."
Right: The model was fixed in this test — and results still swung from 47% to 67% based purely on the harness. A lightweight model in the right harness passed more tasks than most people would expect from any model. The setup mattered more than the raw model.
Wrong: "This is technical benchmark stuff — it doesn't apply to me."
Right: If you're not technical, your choice of harness IS your result — and the winning setup here was the vanilla one a non-technical person would end up with anyway. You don't need to out-engineer anyone. Simple won, on the scoreboard.
Wrong: "AI agents are unreliable, so I'll wait until they're better."
Right: 47% to 67% — same model, same day, same tasks. Some pairings already work far more reliably than others right now. Waiting doesn't improve your setup. Testing does.
Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.
Read the 158-page wins doc →"This model passed X percent" is incomplete without the harness it ran in — the same model can swing 20 points either way. From now on, your first question is: what harness?
Sorting new leads into hot and cold. The first draft of your weekly email. Pulling details into a short summary. Your tasks — not tasks from someone's demo video.
Two harnesses, same model, same instructions — word for word. One variable. Everything else locked.
Did it finish — pass or fail. How long did it take. How much cleanup afterward. An agent that "finishes" but needs twenty minutes of fixing didn't really pass.
A harness update can flip these results next month. The first run takes an afternoon; every re-check takes about an hour.
Here's what task one looks like for real — the exact lead-sort, run through Pi this morning:
Real session, replayed at reading pace: a 6-row leads file sorted into HOT and COLD, written to a file, first call picked — 17.7 seconds.
And task two: rough notes into a 138-word weekly email draft, win up top, call to action at the end — 13.6 seconds. Now imagine scoring these against your current setup.
The first test is one afternoon. Every re-check after is about an hour. You just watched two of the three tasks run in under 18 seconds each — the testing is faster than the guessing.
One model, one set of 30 tasks. Composio was clear: different harnesses suit different models. Pi won for this model. Claude Code was quicker overall. Hermes is built for a much wider range of agentic work. The lesson is: the pairing is the product — test yours.
Here's a builder who fine-tunes small models saying DeepSeek Flash inside the Hermes or Pi harness beat the big-name deep-research tools for his work. Notice the phrasing — he names the model AND the harness together. That's the habit this whole guide is about.
Everyone can run the same models — this test literally used one model for everything. One group knows their pass rate, time per task, cleanup needed. The other group is guessing, and usually getting half the reliability they could be getting without knowing it.
You can wire all of this yourself with the steps above. Or get the whole thing done inside the Agent Operating System — multiple models already integrated, so trying a new pairing takes minutes, not days.
You can't stay wrong — that's the beauty of the loop. Your three-task test takes an hour to re-run, so the moment a better pairing ships, your own scoreboard tells you. Guessers stay wrong for months. Testers are wrong for a week, at most.
The AI didn't get smarter between the best result and the worst result.
The human choices around it changed.
Which harness. Which setup. Whether anyone bothered to measure.
Same model — one setup succeeded two-thirds of the time, another failed more than half the time and ran twice as slow.
The setup was the bottleneck. And the setup is the part you control.
So pick three tasks this week. Run them through two setups. Write down what passes.