Claude Fable 5.1 · token playbook · free

Cut Your Claude Fable 5.1 Tokens By 80% — For Free

Here's how to cut your Claude Fable 5.1 tokens by up to 80%, for free.

Claude Code sends nearly seventy thousand tokens before you type a single word.

Then it sends them again on every single turn.

Three levels of settings decide where your tokens go, and almost nobody has changed them.

I'll also show you the tested results on the popular free token tools, including one that made things worse.

Same plan. Same work. Up to 80% fewer tokens.Your tokens today100%Level 1 · trim the floorLevel 2 · control the growthLevel 3 · pay less per tokenup to −80%Three levels. All free. All settings — no coding.
The actual sources ↓
§1Level zero · the number nobody has seen

69,978 tokens are gone before you type "hi".

What Claude Code sends before your first wordSystem promptalways loadedBuilt-in tool descriptionsalways loadedYour message: "hi"2 tokens69,978 tokens · measuredThat is the floor — and you pay it before your work starts

What you're watching: I ran the same test on my own machine. I typed 2 tokens, and 63,629 tokens loaded before them. Then I ran the same "hi" on Fable 5.1, and it cost $1.11 at list price.

real run · my machine · claude -p "hi" --output-format json
§2It is not once

It sends the whole conversation again on every single turn.

Every reply resends the floor plus everything said so far70kturn 174kturn 281kturn 393kturn 4110kturn 5130kturn 6The floor rides along on every step
§3Where your plan actually goes

Your plan drains by Tuesday — and not on your work.

Where a drained plan wentLoads before your workthe big chunkYour actual workthe small chunkNot on your work — on the stuff that loads before your work
§4Why this matters more on Fable 5.1

Fable 5.1 is Mythos-class. Everything about it is bigger.

1Mtoken context · default
128,000tokens in one reply
always onthinking · you only steer depth
$10 / $50per million · in / out
~2xthe price of Opus 4.8
The Fable 5.1 spec sheet1M windowthe default128K outputin a single replyThinking: ONalways · steer depth$10 in / $50 outabout 2x Opus 4.8A bigger engine burns more when you leave it idling
verified · Anthropic pricing + model docs
§5The multiplier

Same sloppy habits. A lot more expensive.

Same habits, different billBEFORE · 200K MODELwaste × 1smaller windowthinking you could switch offcheaper per tokenNOW · FABLE 5.1waste × a lot1M window fills quietlythinking always onabout 2x the price per tokenVSBigger window + always-on thinking = habits cost more
§6The map for this whole page

Three levels decide where your tokens go.

Three levels, three dialsLevel 1 · The floorin context beforeyou startLevel 2 · The growthhow the sessiongrows as you workLevel 3 · The pricewhat you payper tokenLevel three has one setting almost nobody has changed
Coming up: the free tools, testedsavesHeadroomsaves a bitCavemansavesPonytailcost MOREone of themIf you installed it already, you want to see that part
"Let's start with level one — the floor."
§7Level 1 · the floor keeps growing

Rules, skills and project files pile on every turn.

The same measurement found more, loaded every sessionStartup floor69,978Skills roster · 192 skills13,263Rules folder11,142Project instruction file~2,400Every session. Every turn. Whether you use any of it or not.
the roster is only names + one-line descriptions — not the skill files
§8The biggest invisible pile

88 tools. 40,170 tokens. Just to describe them.

Four connected servers + one browser extension40,170tokens · 88 tool definitions≈ 456 tokens per toolAbout 456 tokens per tool — before it does anything
§9How it sneaks in

Mail, calendar, notes, browser = +40,000 tokens per request.

You connect four handy tools…✉ Maildescriptions📅 Calendardescriptions📝 Notesdescriptions🌐 Browserdescriptions+40,000every request…on a Mythos-class model, before you ask it to do anything
§10Anthropic's fix · tool search

Tool search: up to 85% fewer tool tokens — and better picks.

Anthropic's reported results with tool search onTool tokens · before100%Tool tokens · afterup to −85%Accuracy · Opus 4.5 before79.5%Accuracy · Opus 4.5 after88.1%Accuracy · older model before49%Accuracy · older model after74%Fewer tokens AND better picking

What you're watching: my own machine, the same "hi", one setting flipped. Tool search off loads 89,204 tokens. Tool search on loads 61,591. That is 27,613 tokens of tool descriptions held back.

real run · ENABLE_TOOL_SEARCH=false vs true
§11Why accuracy goes up

Six relevant tools beat two hundred descriptions.

The model searches, then chooses from what mattersChoosing from six beats guessing from two hundred
§12The trap · specific to Fable 5.1

The 5% threshold moves with your window.

Threshold mode: defer only when definitions pass 5% of the window200K window40,170 tokens of tool definitions5% = 10,000DEFERRED ✓ you save1M window · Fable 5.140,170 tokens of tool definitions5% = 50,000ALL LOADED · every turn060,000 tokensBigger window → higher line → nothing gets deferred
§13Same setup, opposite result

Same tools. Bigger window. Nothing deferred.

The exact same setup200K WINDOWsaving ON40,170 > 10,000 linedefinitions held backyou save every turn1M WINDOW · FABLE 5.1saving OFF40,170 < 50,000 lineall 88 definitions loadevery turn · right when tokens cost moreVSThe savings switch turns itself off right when tokens got expensive
THINKING IT? "More room should mean less pressure, right?"

That's what everybody assumes about a 1M model.

A percentage threshold scales up with the window, so the bigger the window, the less gets held back.

§14One more silent switch

Going through a gateway or proxy? Tool search can be silently off.

Tool search shuts itself down on a non-first-party hostYouClaude CodeGateway / proxynon-first-party hostTool search: OFFno error shownjust off+40,000loads every turnNot broken. Not announced. Just off.
§15Level 1 · recap

Most of your usage is decided before you type.

The level-one fix is settings and awareness, not skillKnow your floormeasure it onceTool search: ONnot threshold modeCheck your routegateways can kill itTrim what loadsskills + toolsNone of this is coding — it's knowing which dials exist
§16Want the tested playbook? The exact settings we use

Get the tested Fable 5.1 token playbook — and the Agent OS.

The exact settings we use to cut Fable 5.1 tokens live inside the AI Profit Boardroom — plus the Agent OS setup, where you plug Claude, Hermes and OpenClaw into one dashboard with shared memory, so they stop re-reading the same context over and over.

The Agent OS zip file — Claude, Hermes and OpenClaw in one dashboard with one shared memory
The tested token settings — tool search, pruning, effort and cache — configured the way we run them
A 30-day roadmap — from nothing installed to the whole system running
Four coaching calls a week — share your screen and we'll look at where your tokens are actually going
4,000+ business owners — a lot of them had never used AI before joining
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Link in the description · or join on Skool
"Right. Level two. What happens while you work."
§17Level 2 · how the session grows

A session is not a flat cost. You pay the new total every time.

Each message: you pay for everything so far, againtotalmsg 1total +msg 5total ++msg 10total +++msg 20total ++++msg 40A cost that grows with every message
§18Anthropic's own measurements

Long session: prune −39%. Short session: the results flipped.

LONG agent session · three ways to control growth−39%Manual prune−32%Auto-compaction0%Context editingOn a long run, the prune wins
SHORT session · 20 issues · same levers0%Manual prune0%Auto-compaction+74% costContext editingSame technique. Shorter run. Worse than doing nothing.
that's why blanket advice on this topic falls apart
§19The prune, in one picture

At each task boundary, swap big stale outputs for one line.

The prune that worked bestBEFORE · stale historywhole file that got readlong result that came backanother big tool outputTask boundaryswap for one lineAFTER · what the model needs→ file read: config is valid→ result: 3 tests failed→ output: deploy finishedThe model doesn't need the whole thing — it needs the conclusion
§20The benefit that is easy to miss

The edits sit at the end — so the cache holds.

Measured cache reads around a boundary pruneFirst request after boundary89% from cacheRequests in between81% from cacheRe-sent at full pricevery littleVery little got re-sent at full price
§21The auto-compaction trap

Running to the wall means paying to read everything — 3 or 4 times.

The auto-compaction trap: run to the wall, summarise, repeatfill-up #1read ALL · full rate100,000+ tokens+ pay for the summaryfill-up #2read ALL · full rate100,000+ tokens+ pay for the summaryfill-up #3read ALL · full rate100,000+ tokens+ pay for the summaryfill-up #4read ALL · full rate100,000+ tokens+ pay for the summaryWaiting until you are nearly full is the expensive way to get smaller
§22The habit that matters

New work, new session. Compact at phase changes, not at the wall.

Where to compact, and where to start freshResearchphase ends → compactBuildphase ends → compactUnrelated task?start a new sessionThe wallnever wait for thisCompact at natural phase changes — never at the wall
"Stale context wastes tokens on every subsequent message."— ANTHROPIC · CLAUDE CODE COST DOCS
§23My favourite example · nothing to do with coding

Pasted table: 6 out of 25. Uploaded file: 25 out of 25.

Anthropic's table test · same data, same 25 questionsPASTED INTO THE PROMPT6 / 25~91,000 input tokenson every single requestworse answersUPLOADED AS A FILE25 / 25model runs code against itabout 1/12 of the costevery answer rightVSSame data. Same questions. One twelfth of the cost.
§24Do this today

Customer list? Sales export? Upload it — never paste it.

Where the numbers live changes everything📋 Pasted in chatpaid on every replyworse answers📎 Uploaded filemodel reaches inwhen it needs toBetter answersat a fractionof the costWorth more than every token-saving tool combined
THINKING IT? "I'm not technical — is this one for me?"

This is the most non-technical fix on the page.

If you can attach a file instead of pasting it, you can do this today.

§25A related feature worth knowing

Programmatic tool calling: 24% fewer input tokens, higher score.

Tool results stay in code — only the answer comes backModel writes coderuns the tool callsRaw resultsstay out of contextFiltered resultonly this comes back−24% inputand a HIGHER scoreAnthropic: agentic search tests · fewer tokens, better result
§26The pattern across all of level two

Smaller context is a quality upgrade, not a compromise.

The less raw material you carry, the better it performsworstcarry everythingbettercarry lessbetter stillcarry conclusionsbestcarry only what mattersLess in the window → better answers out
"Now level three. What you pay per token."
§27Level 3 · the big change on Fable 5.1

Cache reads dropped to 25 cents per million.

Fable 5.1 · price per million input tokensFresh input$10.00Cache read · before$1.00 · 0.1xCache read · Fable 5.1$0.25 · 0.025xReading from cache is 40x cheaper than sending it fresh
40xcheaper than fresh input
0.025xbase rate on Fable 5.1
0.1xevery other Claude model
verified · Anthropic pricing docs
§28What a good setup looks like now

Keep the front of your context still.

Same instructions · same order · same tools · same placeFront stays identicalinstructionstoolsrulesnew messagecached · $0.25 / MFront gets shuffled or editedinstructionstoolsrulesnew messagefull price · $10 / MKeep the front of your context still
§29Effort · the fastest way to burn a plan

One step up the ladder: nearly 9x the cost.

Same request · output tokens · cost (published test)lowmedium199 · ~1¢high · default1,764 · ~9¢xhigheven moremaxSame prompt. Same answer needed. Nothing on your screen looks different.

What you're watching: my own run on Fable 5.1 — the same prompt at three effort levels. Low wrote 108 output tokens, high wrote 358, xhigh wrote 963. Same three sentences came back each time.

real run · claude --effort low | high | xhigh · Fable 5.1
§30Anthropic's own guidance

Thinking tokens are billed as output. Start at high.

Thinking is billed as output — and output costs 5x inputRoutine tasksstep DOWNmedium or lowStart here: HIGHthe defaultcheck quality holdsGenuinely hardstep UPxhigh or maxStep down once you have checked quality holds
§31New on Fable 5.1 · in beta

Change effort mid-conversation — and keep your cache warm.

One session · different effort per messagePlaneffort: highThe hard decisioneffort: xhighdeep thinking passTidy upeffort: lowcheap passCachestill warm ✓no rebuild to pay forDeep thinking for the hard decision. Low effort for the rest.
per-message effort control · beta
§32The most expensive habit on this list

Simple task? It doesn't have to be Fable 5.1.

Match the model to the jobRENAME FILES · TIDY A LISTsmaller modelroutine workabout half the pricesame resultHARD, LONG-HORIZON WORKFable 5.1the deep reasoningworth ~2x Opus 4.8when it truly needs itVSA Mythos-class model renaming files is money on fire
"Now — the free tools. The tested results are not what the descriptions promise."
§33The free tools · four popular ones

Four free token tools. Tested, not trusted.

What each one squeezesHeadroomwhat the agent readsRTKshell command outputCavemanwhat the agent saysPonytailwhat the agent buildsFour angles on the same problem — four very different results
§34Tool 1 · Headroom

Headroom: 15–20% on code, 60–95% on JSON — and honest about it.

Headroom · reported savingsCoding agents15–20% fewerJSON / structured data60–95% fewerHoldout control group10% left untouchedA tool that tells you when it cannot prove something
output savings: an estimate with a confidence range — not a claimed number
§35Tool 2 · RTK · pay attention to this one

RTK: −77% on its own benchmark. +7.6% cost on real agent work.

The compression was genuine. The bill still went up.RTK'S OWN BENCHMARK · 73 CASES−77%541,111 → 123,205 tokensmeasures command output(the repo says so upfront)JETBRAINS · REAL AGENT WORK+7.6% costat low reasoning efforthigh effort: no differencetask quality: unchangedVSNot cheaper. More.
§36The whole lesson on this topic

Less text going in does not mean less spend coming out.

What actually happenedOutput shrinksreally · −77%Model sees lesstakes another pathLonger pathmore turns + tokensNet resultmore spendThose paths cost more than the tokens saved
§37Tool 3 · Caveman

Caveman: advertised 65%. Properly paired: 8.5%.

Caveman · the repo's own honest tableAdvertised65%86 coding tasks · output tokens−8.5%86 coding tasks · cost~−10%Short Q&A · output~−50%It can only shorten the parts it controls — the prose
§38Tool 4 · Ponytail

Ponytail: the only real saving — 10.3% cheaper per task.

JetBrains series · 80 paired runs−10.3%cheaper per task46tasks cheaper34tasks MORE expensiveThe direction holds. "It saves you ten percent" is a stretch.
shortens what the agent BUILDS — reuse first, write new last
§39One more thing on reasoning models

On an always-thinking model, terse rules can cost you.

A terseness skill on Fable 5.1THE WRITING SIDEsavesshorter repliesfewer output wordsTHE THINKING SIDEcan costmodel reasons about the rulesthinking is always onbilled as outputVSSave on the writing, pay on the thinking
§40Put all four together

The settings are worth vastly more than the add-ons.

Add-ons vs settings · measuredRTK · add-on+7.6% worseCaveman · add-on8.5%Ponytail · add-on10.3%Headroom · add-on15–20%Boundary prune · setting39%Tool search · settingup to 85%Upload not paste · habit−92% (1/12)Cache read · setting40x cheaperInstalling something feels like progress. Changing a number doesn't.
THE OLD WAY · install add-onslow double digits
  • Headroom: 15–20% on code
  • Caveman: 8.5% on real coding tasks
  • Ponytail: 10.3% — cheaper on 46, dearer on 34
  • RTK: +7.6% MORE at low effort
  • Feels like progress because you installed something
THE NEW WAY · change the settingsup to 80%
  • Tool search hides 40,000 tokens a turn
  • A boundary prune saves 39% on a long session
  • Upload, don't paste: a twelfth of the cost
  • Right effort level avoids an 8x jump
  • Cache reads are 40x cheaper than fresh input
§41The habit that actually matters

It isn't any one tool. It's checking.

Every number on this page came from a measured testRun it WITHthe changeRun it WITHOUTthe changeComparereal numbersKeep or dropyou decide on dataNot from a description — from a test
§42The part that makes this durable

Models change. These four never do.

Fable 5.1 will not be the newest model for longThe flooris always the floorContextalways compoundsCachealways beats freshEffortalways maps to costLearn the levels once — carry them into Hermes, OpenClaw, whatever ships next
§43"I'll wait until it settles down"

The tools rotate every few weeks. The principles have not moved all year.

Waiting is not a planTool Ahot this monthTool Breplaces itTool Creplaces thatThe principlessame all year ✓Know your startup floor. Prune at task boundaries.
Wrong: "I'll wait until the tools settle down, then sort my tokens out."
Right: The tools never settle. The four principles haven't moved all year, so learn those once.
Wrong: "Saving tokens means installing the newest repo."
Right: The measured wins came from settings and habits. The add-ons were low double digits at best.
Wrong: "A smaller context means worse answers."
Right: The uploaded file scored 25 out of 25. The pasted one scored 6. Smaller context was the upgrade.
See what members are doing with this

Members post their wins every day — agency owners, ecom founders, course creators and solo operators across 38 countries.

Read the 158-page wins doc →
§44Your move The full Fable 5.1 token playbook

Get the full token playbook — and the Agent OS it runs on.

The tool search settings, the pruning pattern at task boundaries, the effort levels we use for each type of work, and the cache setup that takes advantage of that forty-times-cheaper read price — it's all inside the AI Profit Boardroom.

The Agent OS zip file — Claude, Hermes and OpenClaw run off one shared memory layer, instead of each one re-reading your context and charging you for it
A 30-day roadmap + step-by-step video walkthrough — for setting the whole thing up
Daily updates — as we improve it
Four coaching calls a week — share your screen and we'll find where your tokens are going
A prompt library built around these workflows — plus a member map so you can connect with people near you
4,000+ business owners — plenty running Fable 5.1 every day, and plenty who'd never touched AI before they joined
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Link in the description · or join on Skool
§45Your homework

Go check your own floor.

What you're watching: the one-line check. Run "hi" with JSON output, add up the three input numbers, and that total is your floor. Mine came back as 63,631. In Claude Code you can also type /context to see what is filling the window.

Look at what loads before you type — and how much you actually use20–30ktokens never neededRules you never opentrimSkills you forgot you installedtrimTools from servers you never usetrimWhat is left: your real floorkeepMost people find twenty or thirty thousand tokens they never needed
"Go check your own floor. See you in the next one."