The Token Killer Stack cuts your Claude Code tokens by 80% — and every piece of it is free.
It's four free tools, and each one kills a different token leak.
This means your AI agents get faster, your limits stop running out, and everything keeps working with the setup you already have.
I tested every single one on my own machine — the receipts are on this page.
One of them cut a single command's output by 92%. I'll show you which one later.
By the end, you'll have the whole stack installed and your tokens will go 5x further.
Here's the thing nobody tells you about Claude Code. Most of your bill isn't Claude being smart. It's Claude being flooded. Giant test logs it didn't need to read. Polite three-paragraph answers. Over-built code nobody asked for. Full price paid for grunt work a free model could do. Four different leaks. So it takes four different tools — and all four are free, open-source repos I've already covered one by one. This guide stacks them.
Run one Claude Code session with an eye on the numbers and you'll see it.
Leak 1 — tool output. Claude runs git diff or your test suite, and thousands of tokens of raw log flood the context. Claude needed twenty of those lines. You paid for all of them.
Leak 2 — chatty replies. "Sure! I'd be happy to help you with that. The issue you're experiencing is likely caused by…" Every filler word is billed at output rates — the most expensive tokens on the menu.
Leak 3 — over-built code. You asked for a cache. You got a factory, an interface, a config system and a debate about edge cases. Every unneeded line costs tokens to write, then costs again every time the file is read back into context.
Leak 4 — full price for grunt work. Renaming variables and writing boilerplate doesn't need a frontier model. But if everything runs through one paid endpoint, you pay frontier prices for janitor work.
The whole stack installs in about 15 minutes, one command each. After that it's automatic — no per-session ritual, no remembering. If you spend more than $20 a month on Claude, it pays for itself the first week.
This is the whole system on one screen. Each layer is free, open-source, and attacks exactly one leak. They don't overlap. They don't fight. They stack.
The biggest leak first.
When Claude Code runs a shell command — git diff, npm test, a file listing — the raw output gets pasted straight into the conversation.
A single messy git diff can be thousands of tokens. Claude usually needs a fraction of it.
RTK (the Rust Token Killer) sits between the agent and the shell. It rewrites commands to their compact form, filters the output, deduplicates the noise, and hands Claude only the signal. It adds about 14 milliseconds. You never see it working — your bill just shrinks.
I didn't trust the README's "90% fewer tokens" claim, so I ran my own battery on my actual repos.
Result: 82.9% saved overall. git diff alone dropped 92%.
And it wired into 14 of my agents — including Hermes — with one hook.
brew install rtk
github.com/rtk-ai/rtk — free, open source, Rust.
Inside the Agent OS: RTK runs under my Free Claude Code engine and the Hermes agents — every shell-heavy build my agents run flows through it automatically. That's why my agents can grind all day without the context flooding.
RTK filters mechanical noise — progress bars, duplicate warnings, unchanged-file spam — not content. In my whole test battery the agent never once failed a task because RTK trimmed too hard. And when output IS the content (like a dirty-file list), RTK passes it through nearly untouched: my worst case was −14%, not −92%.
Leak 2 is the expensive one per token: output. The smartest models charge the most for every word they SAY — and by default they say a lot.
Caveman is a skill file your agent reads at the start of every chat. It says: drop the filler. No "Sure! I'd be happy to help." No three-paragraph warm-ups. Answer like a smart caveman: short, blunt, exactly right.
Three rules keep it safe: code blocks stay byte-for-byte exact, security warnings automatically switch back to full clear sentences, and it never invents fake abbreviations (they tokenize the same anyway).
My test on Fable 5: 69% fewer output tokens, 37% off the total bill, five out of five answers still technically correct.
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash
github.com/JuliusBrussee/caveman — one command installs into 30+ agents.
Inside the Agent OS: Caveman runs as a session hook on my Claude Code — every reply I get while building the OS is caveman-compressed. I keep it on full. My agents grunt, my bill shrinks, nothing breaks.
Leak 3 is sneaky because it looks like value. Ask for a cache, get a custom cache class with a config system. It feels thorough. It's actually a tax — you pay tokens to write it, then pay again every single time that file gets read back into context.
Ponytail makes your agent act like a lazy senior developer — lazy meaning efficient, not careless. Before writing code it climbs a ladder: does this need to exist at all? Does the standard library do it? Does the platform do it? Can it be one line? It stops at the first rung that holds.
The benchmark numbers: 54% less code, ~22% fewer tokens, 100% of tests still passing. And less code today means cheaper context forever — the savings compound every session that touches the file.
/plugin marketplace add DietrichGebert/ponytail
/plugin install ponytail@ponytail
github.com/DietrichGebert/ponytail — two slash commands, MIT-free.
Inside the Agent OS: Ponytail runs alongside Caveman in my sessions — Caveman shrinks the talking, Ponytail shrinks the code. They're the perfect pair: one guards the mouth, one guards the hands.
Ponytail never touches safety: input validation, error handling, security and anything you explicitly asked for get built in full. It only kills the speculative stuff — the abstractions "for later" that later never needs. And every deliberate simplification is marked with a comment naming the upgrade path, so nothing is hidden.
After the first three repos, every request is lean. OmniRoute attacks what's left: who you're paying at all.
It's a local gateway that sits on your machine and routes AI requests across 237 providers — 90+ of them with free tiers. Boilerplate, renames, everyday builds: routed to verified free coding models. $0. The genuinely hard tasks: still your frontier model, now fed lean requests.
And here's the kicker — OmniRoute has RTK and Caveman-style compression built into the gateway itself, squeezing another 15–95% out of every request that passes through, on top of the routing.
npm install -g omniroute
github.com/diegosouzapw/OmniRoute — free, local, one dashboard.
Inside the Agent OS: OmniRoute is a first-class engine in my OS — it runs Hermes on free models, it runs a full $0 coding agent in my OpenCode tab — the one that built three animated demos for zero dollars, and its dashboard at localhost:20128 shows the compression ratio on every single request. Watching tokens die in real time is weirdly satisfying.
Let's be straight about how stacking works, because the cuts hit different parts of the bill.
RTK's −82.9% hits tool output — usually the biggest single chunk of a coding session's context. Caveman's −69% hits the reply tokens — the most expensive ones. Ponytail's −22% hits the code volume, and keeps paying every future session. OmniRoute takes whole categories of work to $0 and compresses whatever's left.
On a shell-heavy session — the kind Claude Code actually runs all day — tool output dominates, so the stack lands in the 80%+ zone. That's where the title comes from: my own RTK battery measured 82.9% on exactly that kind of work. On a pure-chat session with no tools, you'll see less — Caveman's 37% total-bill cut is the honest floor. Either way: same brain, same answers, fraction of the tokens.
The catch is that each number only cuts its own slice — nothing cuts 80% of EVERYTHING. Tool-heavy sessions hit 80%+ because tool output dominates and RTK feasts. Chat-heavy sessions land nearer 40%. Every number above came from my own tests on my own machine — not vendor marketing.
Before
I was burning through Claude Code limits by lunchtime.
Every session started fast, then drowned — test logs flooding the context, three-paragraph answers to one-line questions.
I'd watch a single git diff eat more tokens than the actual fix.
I tried "being careful with prompts". It didn't move the needle, because the waste wasn't in my prompts.
Then I stopped looking for one fix and stacked four.
After
Now the stack runs my whole Agent OS. RTK filters every shell command my agents run. Caveman keeps every reply tight. Ponytail keeps the codebase lean. OmniRoute sends the grunt work to free models. Same models. Same quality of answers. I run agents all day — benchmarks, builds, blog pipelines — without watching a meter. The token anxiety is just… gone.
You can have this too. Same four repos. Same afternoon of setup.
Here's what's happening for the members already running this kind of stack — agency owners, ecom founders, course creators. Different businesses. Same shrinking bills.
Members post their wins as they happen — token bills cut, first agents shipped, whole workflows automated. They're all collected in one doc you can read right now.
Read the member wins doc (158 pages) →You've seen the numbers. Every one of them came from a real test on a real machine.
So here's the deal.
Don't try to install all four tonight. Pick ONE — RTK if your sessions are tool-heavy, Caveman if your bill is mostly replies — and have it running before you sleep. One command. Fifteen minutes, worst case. Because the moment you feel the first layer working, the other three install themselves — you won't be able to leave the savings on the table.
The people sitting still are paying full price for padding. The people who move tonight are the ones comparing bills in the Boardroom next week, grinning.
Be one of those people. One layer. Tonight.
The four repos are the heavy machinery. These are the habits — each one free, each one shaving tokens on top of the stack.
1 · Clear between jobs. Finished a task? /clear. Starting fresh with old context loaded is paying rent on a room you moved out of. Long single task? /compact summarizes the history instead of dragging all of it forward.
2 · Put your CLAUDE.md on a diet. That instructions file rides along with EVERY session. Every stale rule, every dead note in there is billed on every single request. Once a month, delete half of it. Nothing breaks. (I removed a whole retired workflow from mine and saved thousands of tokens per session.)
3 · Give your agent a memory vault. Re-explaining your business every session is the most expensive small talk in the world. A memory system — like my Obsidian second brain — means the agent reads a few tight notes instead of you re-typing context.
4 · Read slices, not files. Ask for "the auth function in server.js", not "look at server.js". Agents can read line ranges — a 40-line slice beats a 2,000-line file. The answer is the same. The bill isn't.
5 · Send the scouts, keep the throne clean. When Claude Code needs to hunt through a big codebase, have it use a subagent for the search. The scout burns its own context on the noisy exploration and reports back one clean summary — your main session only ever sees the conclusion.
6 · Route by difficulty. Rename-this-variable does not need your smartest model. Simple tasks → cheap or free models (Haiku, local, OmniRoute's free tier). Hard architecture → the frontier model. My OS does this as a 3-tier rule and it's the single biggest bill-shaper after RTK.
7 · Plan before you build. One planning pass, then one clean build, beats five build-fix-rebuild loops. Rework is the invisible token tax — every wrong attempt is paid for twice: once to write it, once to read it back and fix it.
8 · Batch your asks. "Fix the header, the footer, and the mobile nav" in one message shares one context load. Three separate messages pay for the same context three times. Group the related work.
The repos cut the waste per request. The habits cut the number of wasteful requests. A lean request you didn't need to send is still the cheapest request of all. Stack both and that's how "more out of Claude" stops being a hope and starts being arithmetic.
The old way: treat tokens like a scarce resource to ration. Shorter prompts. Fewer questions. Anxiety before every big task. Hitting the limit at 2pm and waiting until tomorrow. The bill stays high anyway — because the waste was never in what you asked. It was in the plumbing.
The new way: fix the plumbing once. Four free repos filter the noise, shrink the replies, thin the code, and route the grunt work to free models — automatically, on every request, forever. You ask MORE of your AI, not less. That's the whole point: the goal was never a smaller bill. It's more output per dollar.
Everything in this guide is already installed and configured inside my Agent Operating System. Join the AI Profit Boardroom and you get:
You're not buying a tool. You're getting the operating system I run a seven-figure business on — with coaching calls where we set it up together.
Get the Agent OS →Backwards. Light users feel limits HARDEST — you hit the ceiling mid-task with no backup plan. The stack is one install command per repo. If you can paste four lines this month, you're qualified.
All four are active, open-source projects — and none of them lock you in. Every layer uninstalls as fast as it installed, and your agent works exactly as before. Worst case you lose fifteen minutes. Best case you keep 80% of your tokens. That's not a risky bet.
Some of it, they slowly do — context compaction got better this year. But a vendor will never route your work to their competitors' free models, and never has your exact stack in mind. The plumbing between YOUR tools is always your job. That's exactly why it's an edge.
1 · RTK first (5 min). brew install rtk, run its hook installer, done. Biggest cut, zero behavior change — you'll forget it's there.
2 · Caveman second (5 min). The curl one-liner installs it into every agent on your machine. Try one session on lite if full caveman-speak feels aggressive, then graduate.
3 · Ponytail third (2 min). Two slash commands in Claude Code. From now on, code comes out lean by default — and every future session reading that lean code is cheaper too.
4 · OmniRoute last (10 min). npm install -g omniroute, connect two or three free providers in its dashboard, point your everyday work at it. Watch the compression column while you work — that's the whole stack visible in one number.
5 · Then adopt one habit a week. Start with /clear hygiene and the CLAUDE.md diet. Eight weeks later the whole list from section XI is muscle memory.
Your Claude Code bill has four leaks — flooded tool output, chatty replies, over-built code, and frontier prices on grunt work.
One free repo per leak: RTK filters tool output (−82.9% on my repos). Caveman shrinks replies (−37% off my Fable 5 bill). Ponytail writes less code (−54% code, −22% tokens). OmniRoute routes grunt work to free models ($0) and compresses the rest.
Tool-heavy sessions land 80%+. Chat-heavy sessions land nearer 40%. Every number in this guide came from my own tests.
Then stack the eight free habits — /clear hygiene, CLAUDE.md diet, memory vault, sliced reads, scout subagents, difficulty routing, plan-first, batched asks.
The goal isn't a smaller bill. It's more output per dollar. Ration less. Plumb better. Ask your AI for more.
Everything above — pre-wired, documented, with people to help when something doesn't work first try. That's the Agent OS inside the AI Profit Boardroom: 4,000+ founders, 258 documented wins, 38 countries.
Join the AI Profit Boardroom →