§0 · The Hook
New benchmark — 8 harnesses · 1 model · 30 tasks

The Harness Multiplier + Pi Agent.

Pi Agent plus DeepSeek just beat Claude Code. And Hermes. And almost every big-name agent tool.

Same AI model in every single test. Same 30 tasks.

One tool passed 20 of them. Another passed 15.

Nothing about the AI changed between those runs. Only the tool wrapped around it changed.

By the end, you'll know which setup won, why it won, and the one mistake almost everyone makes when they pick an AI agent.

Same engine. Different car. Completely different drive — that's the whole benchmark.

20/30tasks passed · Pi
4767%same-model swing
$0.028per passed task
132smedian task time
ONE MODEL DeepSeek V4 Flash Pi Agent Claude Code · Codex Hermes · OpenCode Deep Agents · OMP Prime Agent 47–67% task success Same brain. Eight different bodies. A 20-point swing in how often it succeeds.
§1 · What Happened The test

Composio ran one model through 8 different agent tools.

DeepSeek V4 Flash, 30 hard agentic tasks, eight harnesses. The agent has to go do things, use tools, and finish the job on its own.

The sources — click through and check everything ↓
The announcement · what's going on

The scoreboard, straight from the source

Composio builds tooling for AI agents. This thread is round two of their test: same DeepSeek V4 Flash model pushed through Hermes Agent, Pi Agent, Prime Agent and Deep Agents — after Claude Code, Codex, OpenCode and OMP in round one. Pi was the cheapest harness and passed the most tasks.

§2 · The Word That Matters Harness, defined

The model is the engine. The harness is the car.

Claude Code is a car. Hermes is a car. Pi is a car. Same engine, different car, completely different drive.

ENGINE the model THE CAR — the harness → how the model plans → how it uses tools → how it handles mistakes → when it stops the DRIVE One engine. Eight cars. Which car wins?
§3 · The Numbers Pass rates

Pi passed 20 of 30. The rest didn't come close.

Tasks passed — same model, same 30 tasks Pi Agent 20/30 Deep Agents 16/30 Hermes Agent 15/30 Prime Agent 15/24* *6 Prime runs thrown out — the grader literally couldn't score them. More on that trap below.
Thinking it? "Is 20 out of 30 even good?"

These are hard agentic tasks — go do things, use tools, finish alone. Two out of every three, on a flash model built for speed, is the best score on the board. The point isn't perfection. The point is the same model scored 15 in a different wrapper.

§4 · Speed + Cost The same story

The winner was also the fastest and cheapest of the four.

Fair note: in the overall ranking across all eight, Claude Code and OpenCode were quicker than Pi. On pass rate and efficiency together, Pi came out on top.

Median seconds per task Pi 132.2s Hermes 175.5s Deep Ag. 187.1s Prime 242.1s Cost per passed task Pi $0.028 Deep Ag. $0.045 Hermes $0.056 Prime $0.131 The harness that passed the most tasks was also the fastest of these four — and 4.7× cheaper per win than the heaviest one.
§5 · Zoom Out The framework

The Harness Multiplier.

Across all eight harnesses, the exact same model scored anywhere from 47% to 67% task success. The AI never got smarter or dumber. Only the wrapper changed.

One model. A 20-point swing. 47% worst harness 67% best harness       Cost swung $0.019–$0.104 per task. Time swung 122.7–272.4s. Same intelligence underneath.
i.

The Multiplier

The tool you wrap around your AI multiplies the result — up or down. Right harness: more reliable AND faster. Wrong harness: more failures, slower runs, same intelligence.

ii.

The Light Rule

Every extra layer is another place for the agent to get lost. A clean harness gives the model a short path from "here's the task" to "task done."

iii.

The Pair Test

Never judge a model alone. Test the model-harness pair on your own tasks — because the pairing is the product.

§6 · Their Conclusion In their words

One sentence that should change your whole setup.

"The same model delivered 47–67% task success, cost $0.019–$0.104 per task, and took 122.7–272.4s median time per task depending on the harness. Benchmark the model-harness pair you'll actually use, not the model in isolation."

— Composio (@composio), benchmark thread, Aug 10, 2026

§7 · The Shortcut The Harness Multiplier, put to work

Get the Agent OS — and test your own pairings.

The Agent OS plugs your agents — Claude, Hermes, OpenClaw, Free Claude Code — into one system with shared memory. Run the same task through different setups and see which pairing actually works for your business.

The full Agent OS zip — multiple agents, one dashboard, shared memory
A 30-day roadmap to set it up piece by piece
Video tutorials daily, plus updates as we keep improving it
4 weekly coaching calls — bring your model + harness questions live
A room of 4,000+ business owners, plenty running Hermes and testing these exact pairings
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Set up in an afternoon · used in 38 countries · new tools added the week they ship
Thinking it? "Doesn't running an Agent OS burn a fortune in tokens?"

That's the biggest myth about it. The everyday 90% runs on free local models on your own machine, free APIs slot in for more, and for frontier work it drives the CLIs you already pay for — your Claude subscription already includes the Claude Code CLI, and the Agent OS plugs straight into it, so you're never paying twice.

And this benchmark just proved the cheap route can win: the best harness here cost 2.8 cents per finished task.

"The AI never got smarter. The wrapper changed. The results doubled."
§8 · The Vanilla Winner Back to the test

Pi ran with almost nothing added — and won.

A fresh, vanilla install. The only thing bolted on was the MCP server plugin — the standard connector that lets an agent plug into outside tools. No tuning. No special configuration.

What you're watching: the real install from this morning, replayed at reading pace — one npm command, 3 seconds, and Pi 0.84.1 is live with DeepSeek V4 Flash available. The exact winning pair, on my machine.

Then I gave that vanilla install one prompt. Real session, replayed at reading pace: Pi + DeepSeek V4 Flash wrote a 504-line neon snake game in 56 seconds.

And here's that exact game running — real screen recording, played by an autopilot script. Score ticks up, particles fire, game over screen works. One prompt, zero fixes.

Thinking it? "Surely the winner needs some special setup."

Everything above ran on the out-of-the-box install you just watched. That's the point of this whole benchmark: the vanilla setup beat the heavily-engineered ones.

§9 · The Heavyweight The other extreme

Prime Agent choked on its own weight.

Sessions up to 3.5 million tokens and 33 tool calls — like an agent writing itself a to-do list the length of a phone book before starting the job.

PRIME — the heavy build 3.5M tokens in one session 33 tool calls per run 6 runs the grader couldn't score 242s median · $0.131 per win vs PI — the light build vanilla install + 1 plugin 132s median · $0.028 per win most tasks passed: 20/30 The heavy, do-everything setup choked. The light one passed the most tasks in the least time.
§10 · The Flip Old way vs new way

This flips the old thinking on its head.

Old way heavy + slow
  • Grab the biggest model you can get
  • Stack on every plugin and extension
  • Add every fancy layer "just in case"
  • Assume more always equals better
  • Never measure any of it
  • Result: 3.5M-token sessions the grader can't even score
New way light + tested
  • Pick a fast model — flash-class is fine
  • Put it in a clean, light harness
  • Add ONE connector, only if you need it
  • Test the pair on your actual tasks
  • Keep what passes, re-test when versions ship
  • Result: 2 of every 3 hard tasks done, in 132s
The crowd noticed · what's going on

"Pi beating bloated OMP is hilarious"

This reply summed up the whole benchmark in one line. Builders watching the thread weren't shocked that Pi won — they were laughing that the bloated setups lost. The light-beats-heavy pattern is becoming common knowledge.

§11 · Why Light Wins The mechanism

Every extra layer is another place to get lost.

CLEAN HARNESS task done ✓ BLOATED HARNESS task plugin config giant instructions more tools lost… Shorter path, fewer wrong turns. That's why the vanilla install beat the heavyweight.
§12 · Three Beliefs To Drop What's holding you back

The three beliefs that stop people acting on this.

Wrong: "I need the most powerful, top-tier model to get good results."

Right: The model was fixed in this test — and results still swung from 47% to 67% based purely on the harness. A lightweight model in the right harness passed more tasks than most people would expect from any model. The setup mattered more than the raw model.

Wrong: "This is technical benchmark stuff — it doesn't apply to me."

Right: If you're not technical, your choice of harness IS your result — and the winning setup here was the vanilla one a non-technical person would end up with anyway. You don't need to out-engineer anyone. Simple won, on the scoreboard.

Wrong: "AI agents are unreliable, so I'll wait until they're better."

Right: 47% to 67% — same model, same day, same tasks. Some pairings already work far more reliably than others right now. Waiting doesn't improve your setup. Testing does.

Don't take my word for it

Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.

Read the 158-page wins doc →
§13 · The Hidden Lesson Half the story

Every AI leaderboard you've seen is missing a column.

"This model passed X percent" is incomplete without the harness it ran in — the same model can swing 20 points either way. From now on, your first question is: what harness?

"Model X passes 67% of tasks!" model ✓   tasks ✓   score ✓ harness … missing WHAT HARNESS? the missing column That answer might matter more than the model name at the top of the chart.
§14 · Do This Your own mini-benchmark

Composio used 30 tasks. You need three.

Write down three real tasks from your actual week.

Sorting new leads into hot and cold. The first draft of your weekly email. Pulling details into a short summary. Your tasks — not tasks from someone's demo video.

Run those same three tasks through two setups.

Two harnesses, same model, same instructions — word for word. One variable. Everything else locked.

Score them the same way Composio did.

Did it finish — pass or fail. How long did it take. How much cleanup afterward. An agent that "finishes" but needs twenty minutes of fixing didn't really pass.

Keep the winner. Re-test when new versions ship.

A harness update can flip these results next month. The first run takes an afternoon; every re-check takes about an hour.

Here's what task one looks like for real — the exact lead-sort, run through Pi this morning:

Real session, replayed at reading pace: a 6-row leads file sorted into HOT and COLD, written to a file, first call picked — 17.7 seconds.

And task two: rough notes into a 138-word weekly email draft, win up top, call to action at the end — 13.6 seconds. Now imagine scoring these against your current setup.

3 real tasks from YOUR week 2 harnesses same model + prompt score 3 columns pass · time · cleanup keep the winner re-test on updates the loop — an hour, every time a new version ships
Thinking it? "I don't have time to benchmark anything."

The first test is one afternoon. Every re-check after is about an hour. You just watched two of the three tasks run in under 18 seconds each — the testing is faster than the guessing.

§15 · The Honest Note Don't oversell it

This does not mean "everyone switch to Pi."

One model, one set of 30 tasks. Composio was clear: different harnesses suit different models. Pi won for this model. Claude Code was quicker overall. Hermes is built for a much wider range of agentic work. The lesson is: the pairing is the product — test yours.

Proof in the wild · what's going on

Operators are already pairing this exact model

Here's a builder who fine-tunes small models saying DeepSeek Flash inside the Hermes or Pi harness beat the big-name deep-research tools for his work. Notice the phrasing — he names the model AND the harness together. That's the habit this whole guide is about.

"The pairing is the product."
§16 · The Real Gap Testers vs guessers

The gap isn't access to AI. It's who measures.

Everyone can run the same models — this test literally used one model for everything. One group knows their pass rate, time per task, cleanup needed. The other group is guessing, and usually getting half the reliability they could be getting without knowing it.

§17 · Your Move Be in the group that knows

Get the Harness Multiplier built for you.

You can wire all of this yourself with the steps above. Or get the whole thing done inside the Agent Operating System — multiple models already integrated, so trying a new pairing takes minutes, not days.

Daily tutorials — step-by-step videos, including harness setups like Hermes and running models like DeepSeek through them
Model swapping inside the Agent OS — test pairings on your own business tasks, just like this benchmark did
4 weekly coaching calls — bring your exact setup: your model, your harness, your tasks, reviewed live
A 30-day roadmap so you're not guessing what to do first
The prompt library — lead follow-up, content drafts, customer replies
A member map + 4,000+ members · support around the clock
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Set up in an afternoon · used in 38 countries · someone's always online
Thinking it? "What if I test and still pick wrong?"

You can't stay wrong — that's the beauty of the loop. Your three-task test takes an hour to re-run, so the moment a better pairing ships, your own scoreboard tells you. Guessers stay wrong for months. Testers are wrong for a week, at most.

§18 · The Close The truth this proved

The intelligence was never the bottleneck.

The AI didn't get smarter between the best result and the worst result.

The human choices around it changed.

Which harness. Which setup. Whether anyone bothered to measure.

Same model — one setup succeeded two-thirds of the time, another failed more than half the time and ran twice as slow.

The setup was the bottleneck. And the setup is the part you control.

So pick three tasks this week. Run them through two setups. Write down what passes.

"The answer was never just the model at all."