Open source · Apache 2.0 · tested on my own Mac

Needle 3 AI Runs Offline In 66ms

The whole model is 35 megabytes. Smaller than a song.

Needle 3 is a new AI model that runs in 66 milliseconds with no internet and no graphics card, and the whole thing is just 35 megabytes.

That makes it smaller than most of the apps on your phone, and it can still turn a messy spoken sentence into exact actions.

I installed it on my own Mac and ran every test you're about to see, including the four places where it breaks.

Stick with me, because the part that changes how you think about AI agents has nothing to do with light switches.

A model small enough for a watch, doing a job we used to send to a giant data centre.

66 ms
three actions, on a CPU
35 MB
the whole model
11,000+
GitHub stars
Free
Apache 2.0 · open source
What Needle 3 does, in one pictureA messy sentencesaid the way anormal person talksNeedle 335 MB · on a CPUno internet Exact actionsfan on · temp 10bedroom light on
Official sources — read it and try it yourself ↓
"…an empty list, not a guess."— Cactus Compute, Needle 3 readme, on what happens when no tool fits
§1The live test

One messy sentence in. Three exact actions out.

What you're watching: I typed the exact messy sentence from the test, and Needle 3 returned fan on, temperature 10 and bedroom light on in 68 milliseconds on my Mac.

real run on my Mac · replayed at reading pace
THINKING IT? "This just sounds like a smart home gadget."

The light switch is only the demo.

The same trick pulls the fields out of an invoice or an enquiry, and that's the part your business can use.

§2The speed

Sixty-six milliseconds. On a CPU.

How long one answer takesCloud AI reply~2,000 msA human blink100 msCactus: 2 calls, no GPUunder 97 msMy Mac: 3 actions68 msMy Mac: 2 actions51 ms68 and 51 ms measured on my Mac · 97 ms is Cactus's CPU figure
No graphics cardNo internetNo server farmJust a normal CPU
§3The size

Smaller than the apps on your phone

How big the model file is7B model at 4-bit~4 GBNeedle 3 + engine35 MBNeedle 3, largest file29 MBNeedle 3, smallest file8 MB35,335,380 bytes on my disk · small enough to sit inside a watch
Apache 2.0FreeOpen source11,000+ GitHub stars
§4What it actually does

One file. Three jobs. All offline.

One 35 MB file, three jobs1 · Tool callspicks the function, fills the blanks2 · Structured extractionmessy text in, clean fields out3 · Embeddingssearch and match on the device One fileneedle3.cact35 MB · offline
§5Job one

Tool calls: right tools, right order, or nothing

What you're watching: two requests in one sentence come back as two calls in order, and a question no tool covers comes back as an empty list instead of a made-up answer.

real run on my Mac · replayed at reading pace
§6Job two

Structured extraction: messy text in, clean fields out

What you're watching: I declared vendor, total and due date, handed it an invoice description, and got three clean fields back in 75 milliseconds.

real run on my Mac · replayed at reading pace
§7Job three

Embeddings: search and match without sending anything

What you're watching: eight tools turned into numbers once, then each question matched on the device in under 3 milliseconds. It got 4 of 5 right in my test, and I left the miss in.

real run on my Mac · replayed at reading pace
§8The honest trade

No chat. On purpose.

The trade Cactus made on purposeWhat they gave upNo chat✕ general chat✕ trivia like capital cities✕ essays, poems, small talkWhat they got back121M params✓ beats models 10x its size✓ matches 2-3x bigger on extraction✓ fits on a watch
It's not broken. That's the design.
§9The benchmarks

Six test suites. Thousands of rows.

What Cactus tested it on961Mobile Actions · exact200DroidCall · calls in order3,641BFCL v4 · must refuse too1,813DSTC8 · dialogue turns700SNIPS gold · schema given700SNIPS 7-way · pick 1 of 7
Cactus's published numbers
§10The handicap

The tiny one was handicapped and still won its category

Mobile Actions: phone commands, exact match %DeepSeek V4 Flash88.4Needle 3 · 121M · 2-bit86.0LFM2.5 · 1.2B82.4Qwen3.5 · 0.8B76.0FunctionGemma · 270M65.1Apple FM · 3.0B57.6Needle ran as its shipped 2-bit file · the others ran at full precision under vLLM
Cactus's published numbers
DroidCall: 47.0 — top on-device modelExtraction: level with 2–3x bigger modelsBFCL v4: 50.2 — mid-pack, not a win
§11On a Raspberry Pi

A computer that costs less than a nice dinner

Tokens per secondPi 5 · input, up to10,000/sPi 5 · output, up to4,000/sMy Mac · input2,720/sMy Mac · output1,256/sPi 5 range: 400–4,000 out, 1,000–10,000 in (Cactus) · Mac numbers from my runs
§12Quick pause

"Cool, but I'm not a developer"

The shift changes three decisions this yearWhat you buildsmall jobs firstWhat you pay forthinking only What you stoppaying for
THINKING IT? "I'm not a developer, what do I do with this?"

You're not going to compile this yourself, and you don't need to.

You need to understand the shift, because it changes what you build and what you pay for this year.

§13Inside the AI Profit Boardroom Over 3,000 business owners · many had never touched AI before

Let small free models do the boring jobs — and save your big agents for the thinking.

This is what we break down on the coaching calls: where small local models like Needle 3 take over the lookup-and-fill jobs inside your automations, while your bigger agents handle the thinking.

Four coaching calls every week — bring your actual setup and get it sorted live
The Agent OS — plug your Claude, your Hermes, your OpenClaw into one dashboard with one shared memory
The zip file — the same system I run, ready to install
The 30-day roadmap — to set it up in your business, step by step
The video walkthrough — so you can follow along click by click
Daily updates — every time we improve it, you get the new version
Join the AI Profit Boardroom → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Link in the description · used in 38 countries

What you're watching: the real Agent OS dashboard, where every agent plugs into one place and shares one memory.

THINKING IT? "Doesn't running an Agent OS burn a fortune in tokens?"

No. The everyday work runs on free local models and free APIs, and the frontier work drives the CLIs you already pay for, like the Claude CLI inside your Claude subscription.

Inside the Boardroom there are token-efficiency tutorials too, so you learn to cut usage and stop thinking about it.

§14The bigger story

Needle 3 isn't alone. Meet Jev.

Two cousins, one ideaJev · TypeSafe AI~70 ms• only outputs structure• bigger model, runs in the cloud• trained hard on classification• plays Doom, ~10 moves a secondNeedle 3 · Cactus68 ms• only outputs structure• 8–35 MB, runs on the device• open source, Apache 2.0• free to run, works offline
§15The shift

Models that don't talk. They act.

No paragraph. No preamble. Straight to the action.A requestplain wordsStructured outputno chatty preamble An actionnothing to strip out
§16The problem

The 2am JSON Problem

The layer you didn't know you were paying forAgent writes a replyone word at a timeFind the JSON in itsomewhere in thereCheck it isn't mangleda sentence snuck in? Turn it into an actionbreaks at 2am
OLD WAY
~2 seconds · metered · breaks sometimes
  • A giant model writes a reply one word at a time
  • Something has to dig the JSON out of the reply
  • One extra sentence before the JSON breaks the workflow
  • Every call goes over the internet and gets billed
  • You find out it broke at 2am
NEW WAY
~66 ms · free to run · always parses
  • A 35 MB model outputs the structure directly
  • Every token has to fit the shape you asked for
  • It can't invent a field that doesn't exist
  • It runs on the device with no internet
  • The big expensive model only gets called for thinking
THINKING IT? "My automations work fine most of the time."

Most of the time is the problem, because the one broken run is the one a customer sees.

This is a reliability gain more than a speed gain, and reliability is what makes automations survive past week two.

§17Why it can't break

Every token has to fit your shape

A grammar built from your own schemaYour schemavendor · total · dueByte grammarbuilt from schema Every token fitsso it always parses
Can't hallucinate a fieldCan't hand you broken dataA reliability gain, not a speed gain
§18The cleverest design choice

Intelligence laddering: 19 models in one file

Every depth from 2 layers to 20 is its own working model25M2 layers29M4 layers52M8 layers98M16 layers121M20 layerspick the rung your hardware can carry · the full thing is the top step

What you're watching: I built three rungs on my Mac with one command each, and the 2-layer rung came out at 13.34 megabytes.

real run on my Mac · replayed at reading pace
§19Where the parameters live

Most of it is memory, not maths

The engram: a lookup memory of common word patternsParameters it stores121MSitting in the engram70.8MCompute it really does~50Mover 2x fewer operations per token than a normal transformer of the same shape (Cactus)
Cactus's published numbers
§20Fine-tuning

Tune any rung. 29 million parameters beat a cloud model.

DroidCall accuracy % — before and after one narrow fine-tune4 layers · 29M · before34.54 layers · 29M · tuned62.5DeepSeek V4 Flash60.520 layers · before52.020 layers · tuned70.0every rung gained 18 to 36 points · trained for one epoch on one narrow task
Cactus's published numbers
§21What this means for a normal business

Narrow the job enough, and small wins

What you're watching: an example enquiry goes in, and the name, company and budget come out in 71 milliseconds with no API call.

real run on my Mac · replayed at reading pace
THINKING IT? "I don't have a big dataset to train anything."

You don't need one, because the job is narrow.

A few hundred examples of your own enquiries is the whole dataset.

Narrow the job enough, and small wins.
§22How the fine-tune works

The fine-tune is two commands

Examples in, your own tiny model outYour examplesdata.jsonlAn adapterhow many passes? Your tuned filepick platform + layers
pip install "cactus-needle[train]" needle finetune data.jsonl --epochs 10 --out adapter.safetensors needle build --lora adapter.safetensors --layers 8 --out tuned.cact
§23The part to pay attention to

Every answer comes with a confidence score

Three buckets you can hang your automations onHigh → run itmy test: 1.0 → fan, temp, lightMiddling → show it and askmy test: 0.63 → 'switch that on'Too low → held backmy test: 0.28 → nothing reached me Confidencea trained scoring headon every response
Engine floor: 0.1Not a probability guess — a calibrated numberMost setups run everything or nothing
§24Triggers

Intents that must never be missed

What you're watching: the same vague requests scored 0.63 and 0.70 without triggers, then 1.0 and 0.89 once I attached the words turn, switch, power and flip to the tool.

real run on my Mac · replayed at reading pace
§25The honest part

Four real problems when people pushed it

Where it breaks matters more than where it shines1 · Historyleaks without a reset2 · Defaultsmatter enormously3 · Messy texthas limits4 · Classificationnot its strongest job
THINKING IT? "If it breaks, why would I bother with it?"

Because every one of these four has a simple fix, and I show each one below.

A specific tool with known edges is easier to trust than a big one that fails at random.

§26Problem one

History leaks — so reset between requests

What you're watching: office to 21 degrees, then a pizza, then the capital of France. On engine 3.0.2 no pizza call was invented, but confidence dipped without a reset, so I reset anyway.

real run on my Mac · replayed at reading pace
Reported by testers: an invented pizza callCactus shipped an engine updateMy run on 3.0.2: zero invented calls
§27Problem two

Give every argument a sensible default

What you're watching: the same booking request was held back at 0.28 confidence with no defaults, then ran at 0.95 the moment I added them.

real run on my Mac · replayed at reading pace
§28Problem three

Messy text has limits

Same invoice, two ways of writing itA clean description3 of 3 fields✓ vendor → Brightline Studio✓ total → 2450.0✓ due → 2026-10-15A rambling chat message0 of 3 fields✕ vendor → empty✕ total → empty✕ due date → empty
real run on my Mac · replayed at reading pace
My extra finding: "$1,500" came back as 1200"1500 dollars" came back rightThe shape of what you feed it matters
§29Problem four

It's weaker at pure classification

If sorting things into categories is your main jobNeedle 3Can do it• can do it, using enums• not what it trained hardest on• my test: picked the wrong serviceJevBuilt for it• trained for exactly this• a bigger model• the stronger pick for sorting
Specific tools are the ones that actually hold up.
§30Where it runs

A watch. A television. A car. A browser tab.

Every platform ships a prebuilt enginemacOSApple siliconLinuxx86·ARM·ARMv7·RISC-V·MIPSWindowsx86 · ARMAndroidthree chip typesiOSiPhone + iPadtvOSthe televisionwatchOSon your wristBrowserWebAssembly, no serverWASIa portable component
Under 1 MB of engine, then it loads the weightsA watchA televisionA carA robot vacuumA browser tab, no server The engineunder 1 MB368 KB on my Mac
§31One practical note

Telemetry is on by default

Turn it off with two settings2 variablesCactus documented this themselves in the readme
export NEEDLE_TELEMETRY=0 export DO_NOT_TRACK=1 # I ran every test in this guide with both set
§32What this means going forward

For two years the direction was bigger

Two directionsThe last two yearsBigger✕ more parameters✕ more compute✕ more cost per call✕ everything through an API✕ everything meteredWhere this pointsRight-sized✓ read this, pick that, fill these✓ runs where the person is✓ instant✓ free to run✓ works on a plane
§33The investment

360 billion tokens of one kind of thing

What Cactus trained it on360B tokensnot knowing more things — knowing one kind of thing properlyMost of what your business needs from AI isn't intelligence.It's the same small judgement, made correctly, a thousand times a day.
§34The takeaway

The right size model for the right job

What a routing layer doesHeavy agents → the thinkingClaude · Hermes · OpenClawSmall model → fetch and fillNeedle-style · instant · free to runOne shared memorySEO agents · avatars · 1 dashboard Routing layerdecides whatgoes where
how the pieces fit · concept diagram

What you're watching: the real Agent OS, where every agent reads from the same shared memory.

§35Day to day

Every enquiry, filed before anyone reads it

No API call. Nothing leaves the machine.An enquiry landsemail · form · DMDetails pulled71 ms in my test Filedbefore anyone reads it
§36If you're thinking this is too technical

Three beliefs to drop

Wrong: "Bigger models are always better."
Right: A 121 million parameter model scored 86.0 on phone commands and a 1.2 billion one scored 82.4, because its job was narrow.
Wrong: "Every AI job has to go through a paid API."
Right: Read this, pick that, fill these fields can run on the device, and the expensive model only gets called when it's needed.
Wrong: "This is all too technical for me."
Right: The install is one line. The concepts took me longer, and you now have them.
Plenty of members had never used AI at all

Members post their wins every day — agency owners, ecom founders, course creators and solo operators across 38 countries, in their own words.

Read the 158-page wins doc →
§37If you want help actually building this Over 3,000 business owners · someone's always online

Get every agent running in one place — and get it live in your business in 30 days.

The Agent OS where all your agents run in one place, the routing so small models handle the cheap jobs and your big agents handle the rest, and a roadmap that gets it live.

Four coaching calls every week — bring your setup and we fix it live
Daily tutorials — how these releases get used to save time and bring in more customers
The Agent OS zip file — plus the video walkthrough
Daily updates — every time we improve it
A 30-day roadmap — that gets it live in your business
A prompt library — ready to copy
A member map — find people near you doing the same thing
Join the AI Profit Boardroom → Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Link in the description · 4 live calls a week · used in 38 countries
§38Last thing

Two years ago that was a research paper. Now it's a pip install.

The whole thing, in three facts35 MBthe whole modelRaspberry Piit runs on oneApache 2.0use · change · ship it
pip install cactus-needle

What you're watching: the real one-line install on my Mac, the 35 megabyte download, and the model loading with no GPU and no API key.

real run on my Mac · replayed at reading pace
The right size model for the right job. The expensive one only when it's needed.