Google just fixed the one thing that made local AI agents unreliable. Download one file, wire one profile, and your agent works on your own machine — $0 per message, private, and it never expires.

Imagine hiring a brilliant assistant.
They're great at the job.
But they don't work for you — they work for a landlord in another city.
Every sentence they speak, the landlord charges you.
Every file they touch gets photocopied and sent to the landlord's office.
If the landlord raises prices, you pay.
If the landlord changes the rules, revokes your key, or has an outage — your assistant vanishes mid-task.
And if your internet drops, they stop existing entirely.
That's every cloud AI agent you run today. A genius on a meter, in someone else's building.
The Self-Owned Agent breaks that cycle for good.
It's three commands, copy-pasted. One app install, one download, one text file.
If you can install Spotify, you can do this. The full setup is in this guide.
Three pieces, all free.
Gemma 4 is Google's open AI brain — a model file anyone can download and keep. On July 15, 2026, Google shipped an update that patched the one weakness local models always had: they were clumsy at using tools (the ability to actually run commands and edit files, not just talk). That patch is what makes this whole setup work now.
Ollama is a free app that runs AI brains on your own computer — think of it as the engine room. It downloads the brain once and serves it to anything on your machine.
Hermes is the agent shell — it gives the brain hands. It's what turns "a chatbot" into "an employee that creates files, runs commands and finishes jobs."
Wire the three together and you own the whole stack. No account, no key, no meter, no landlord.
That was true last year. The July update specifically fixed tool-calling — the exact skill an agent needs.
I tested it the day I set this up: it called tools correctly on the first try, wrote a real file, and verified its own work. The receipts are below.
The brain is literally one file. Gemma 4 (12B) is a 7.7GB file sitting in a folder on your Mac (~/.ollama/models — "the AI brain shelf"). Once it's downloaded, the internet is optional.
Ollama serves it like a tiny local website. The engine app listens at an address that only exists inside your computer (localhost — "this machine, talking to itself"). Nothing on that address can be seen from outside.
You type a message to Hermes. Hermes bundles your message and sends it to that local address — not to Google, not to any cloud.
The brain thinks on your own chip. Your Mac's processor runs the model. On my machine a reply starts in about 1.5 seconds.
When the job needs hands, the brain asks for tools. "Tool-calling" means the model replies with a structured request — "run this command", "write this file" — instead of just words. This is the part Google's July update fixed. Hermes executes it and hands the result back.
The loop repeats until the job is done. Think → act → check → continue. That's what makes it an agent, and every loop costs the same: nothing.
Nothing is metered, logged, or sent away. No token counter, no usage dashboard, no key to revoke. The bill for 10,000 messages is identical to the bill for one: $0.
So when someone asks "but what IS it?" — it's a brain file on your disk, an engine app that runs it, and an agent shell that lets it do real work. All three are free.
On July 15, 2026, the Gemma team announced a big improvement release for Gemma 4 — and buried in the thread is the line that matters for agents: patched chat templates and tool-calling fixes for "accurate, consistent tool execution."
This is the post that triggered this whole guide. The improvements roll out across the Gemma 4 family — and the tool-calling patch is the one that turns Gemma from "a chatbot you run locally" into "an employee you run locally." That's the unlock.
The family, in plain words: five sizes. E2B and E4B are the small ones for lighter machines (128K context — "context" is the AI's working memory). 12B is the sweet spot for a normal Mac. 26B and 31B are the big ones for heavy hardware (256K-class memory). This guide uses the 12B.
How I tested, in plain English: my testbed is my own Mac (36GB of memory) running my real Agent OS dashboard. Every test is one command or one message, sent the normal way I'd use it in a working day. I measured the reply times the tools reported. Nothing cherry-picked — the numbers below are the only runs I did, first attempt each.
gemma4:12b-mlx — the version tuned for Apple chips) and asked it a simple question directly.
What you're looking at: my actual Agent OS chat, talking to the new Gemma profile. It answers as the local model — this reply cost $0 and never left the machine.
Keep it. The point isn't replacing your best model — it's stopping the meter on the everyday 90%: drafts, summaries, file jobs, research grunt work.
Route the routine work to the free local brain, save the paid calls for the hard 10%. That's how the whole Agent OS runs.
This is the repeatable method. Set up each layer once, own the whole stack for good.
Download the brain — Gemma 4, one 7.7GB file from Google, free and yours to keep. This is the layer nobody can take back.
Install Ollama so the brain runs on your own chip. One free app; it serves the model to everything else on your machine.
Wire it into Hermes with one profile. Now it doesn't just talk — it writes files, runs commands and finishes real jobs.
Surface it in your Agent OS dashboard next to every other agent you run — one screen, and the free one handles the daily grind.
Enjoy what ownership actually means: no meter, no key, no logs leaving the building, no landlord. The agent is property now, not a subscription.
You can wire this yourself with the steps below. Or get the whole thing done inside the Agent Operating System — the free local brain, the Hermes profile and the dashboard already connected.
Step 1 — Install Ollama. Grab it from ollama.com/download and open it once. That's the engine room done.
Step 2 — Pull the brain. Open the Terminal app (the text window where you type commands) and paste:
ollama pull gemma4:12b-mlx
That downloads the 7.7GB July-2026 build tuned for Apple chips. On a smaller machine (8GB), use gemma4:e4b instead.
Step 3 — Sanity-check it.
ollama run gemma4:12b-mlx "say hello in five words"
A reply means the brain is alive on your chip.
Step 4 — Give it hands with a Hermes profile. Create a folder ~/.hermes/profiles/gemma4 and drop in a config.yaml pointing at the local engine:
model:
default: gemma4:12b-mlx
provider: ollama-launch
base_url: http://127.0.0.1:11434/v1
context_length: 65536
ollama_num_ctx: 65536
providers:
ollama-launch:
api: http://127.0.0.1:11434/v1
default_model: gemma4:12b-mlx
models:
- gemma4:12b-mlx
name: Ollama (local, offline)
toolsets:
- hermes-cli
terminal:
backend: localStep 5 — Run your first free agent job.
hermes --profile gemma4 -z "Create hello.txt saying: Gemma 4 runs Hermes free forever. Then read it back and confirm."
Mine wrote the file, read it back, and confirmed — the full agent loop, $0.
Step 6 — See it in your Agent OS. The dashboard picks the profile up automatically — it appears in the Hermes tab next to your paid agents, ready to chat.

What you're looking at: the Hermes command centre inside my Agent OS — the live agent your new profile plugs into (367 sessions running through it on this machine). The moment the gemma4 folder exists, it appears as a profile pill in the chat — no restart needed.
The 12B brain needs roughly 8GB of memory while running — most Macs from the last few years handle it.
On an 8GB machine, use the E4B size instead: same July fixes, smaller brain, still free.
That's the biggest myth about it. The everyday 90% runs on exactly this — a free local model, $0, nothing leaving the machine.
Free API tiers plug in on top, and for frontier work it drives the CLIs you already pay for — your Claude subscription includes the Claude Code CLI, so you're not paying twice. Plus the Boardroom has full token-optimisation tutorials.
Wrong: "Serious AI has to live in the cloud."
Right: Google now ships its open brain as a file you keep. The cloud is a convenience, not a requirement — and for routine agent work, the local version is already good enough. I verified it the day the update landed.
Wrong: "Free means it'll cost me in quality."
Right: Free here means no meter — the model itself is the same one Google publishes for everyone. The July tool-calling patch was the missing piece, and it passed every agent test I threw at it, first attempt.
Wrong: "I'll set this up later, when local AI matures."
Right: It matured this month — that's what the update was. The people wiring free local brains into their stack now are the ones whose costs flatline while everyone else's climb.
Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.
Read the 158-page wins doc →You stopped paying per thought. The meter is gone — 10,000 agent messages cost what one costs: nothing.
You own the brain outright. A 7.7GB file on your own disk that no company can revoke, reprice or retire.
Your work stays home. Prompts, files and client data never leave your machine — private by physics, not by policy.
Your agent has real hands now. The July update fixed tool-calling, so the free brain writes files and runs commands reliably — verified, not hoped.
Your stack survives anything. Outages, rate limits, price hikes, plane mode — the Self-Owned Agent doesn't notice.
Should you set it up this week? My advice: pull the model tonight — it's one command and the download runs while you sleep. The people who figure out AI agents now, while the tools are evolving fast, are going to be way ahead when everything settles. Every workflow you move off the meter compounds.
This guide gives you the Self-Owned Agent — one free brain, wired and working. The Boardroom gives you the whole operating system around it: the memory that knows your business, the kanban agents that ship while you sleep, every CLI you already pay for on one dashboard, and the playbooks that make all of it cheap to run.
Readers bookmark this page and keep paying the meter. Operators join, install the Agent OS this week, and move their routine work to $0 by Friday.
Decide which one you are tonight.
Get the Agent OS → Inside the AI Profit Boardroom · skool.com/ai-profit-lab