Gemma 4 just got fast on Apple Silicon — fast enough to run real agentic loops on your own Mac, for free. Not "build a game" — actual working workflows: triage, research, self-critique, a 24/7 watchdog, an overnight agent crew. Here are the six I run with Hermes, a tuned speed profile, and real runs you can copy.
This is the reason agentic loops on a local model suddenly make sense. Ollama turned on multi-token prediction for Gemma 4 via Apple's MLX engine. Agentic workflows loop many times — read, think, call a tool, repeat — so speed matters enormously. A local model that's now this fast (and free) can run those loops all day without a meter.
Before
I'd set an AI agent on a task and then sit there watching it, because every loop cost money on a cloud model.
So I kept the loops short and shallow — no real iterating, no running things overnight.
Anything "agentic" felt expensive, so I mostly used AI like a fancy chatbot.
The good stuff — agents that loop, check their own work, run all night — stayed theoretical for me.
Then Gemma 4 got fast enough locally that a loop costs nothing but a bit of my Mac.
After
Now I set a loop going and walk away — it can run all day, free.
My agents triage, research, critique their own drafts, watch things for me overnight.
I stopped counting tokens and started running real agentic workflows.
Same idea you've heard about "AI agents" — finally cheap enough to actually run.
I run an AI agency with 70+ people where AI handles about 80% of the ops, and a room of operators automating with agents across every kind of business.
No invented quotes. The wins are real and written by the members themselves — agency owners, ecom founders, creators, solo operators across 38 countries. Read them in their own words.
Read the 158-page wins doc →Below are six agentic workflows, three of them with real runs you can copy line-for-line.
Here's the deal I want to make with you.
Pick ONE workflow tonight — the triage loop, the draft-critique-revise loop, whatever fits your day — set up the speed profile, and actually run it once. Free, on your own machine.
Because the moment you have an agent that loops on its own for , you stop using AI like a chatbot and start using it like a worker. That's the whole shift.
Be one of the people who runs it today.
Commit to the transition. One agentic loop, running tonight.
Every workflow below runs on one tuned Hermes profile — Gemma 4 on the fast MLX engine, set up for long agentic loops. Drop this in as ~/.hermes/profiles/gemma-speed/config.yaml and start it with hermes chat --profile gemma-speed:
model:
default: gemma4-mlx:latest # fast MLX+MTP build (~62 tok/s)
provider: ollama
base_url: http://localhost:11434/v1
toolsets:
- hermes-cli # let it use shell + tools
agent:
max_turns: 200 # long loops are cheap here — let it run
gateway_timeout: 1800
Two tuning notes: max_turns: 200 lets an agentic loop run long without you re-prompting, and set Ollama's keep_alive: -1 so the model never unloads between calls. Fast model + long leash + zero cost = loops you'd never dare run on a paid API.
These aren't demos of "build a game." They're jobs an agent does for you, on its own, in a loop. Three of them below have a real run I captured on the gemma-speed profile.
The agent writes something, then critiques its own work against a standard, then rewrites it — looping until it's good. This is the loop a fast free model is perfect for, because "loop 5 more times" costs nothing. Great for headlines, cold emails, product copy, ad hooks.
Point the agent at a file, a folder, a log — it uses its shell tool to read it, then hands you a clean summary. Because it runs locally, you can point it at private files without anything leaving your machine.
The agent reads your new messages, sorts them (urgent / reply / ignore), and drafts replies to the easy ones — leaving you a short list of what actually needs you. Runs on your own machine, so your inbox stays private. Free means you can run it every hour.
A loop (on a cron) that watches something — a page, a folder, a number — and only pings you when something actually changes or breaks. Because it's free and local, leaving it running all day costs nothing. This is the one that quietly earns its keep.
Drop a pile of tasks on a Hermes kanban board and let several Gemma 4 workers churn through them in the background — research five competitors, draft ten replies, tidy a folder. You wake up to finished work, not a running meter.
Talk out a messy brain-dump; the agent transcribes it, pulls the real action items, creates tasks, and drafts the follow-ups. A loop that turns two rambly minutes into a clean to-do list. Local, so your voice notes stay yours.
The draft → critique → revise loop is the clearest example. On a paid model you'd stop after one or two passes to save money. On a free local model, you let it spin until it's genuinely good.
Every loop above — triage, research, self-critique, the watchdog, the overnight crew — is wired into one dashboard in the Agent OS, sharing one memory and running on the fast free local model.
You're not buying a tool. You're getting the whole operating system I run a seven-figure business on.
Get the Agent OS →No — and these agentic loops are exactly why it doesn't. The Agent OS runs the everyday 90% on a free local model — now a fast one — on your own machine, so loops that would rack up a cloud bill cost $0 and never leave your Mac. Free APIs slot in for more, and for the frontier stuff it drives the CLIs you already pay for (your Claude subscription already includes the Claude CLI — the Agent OS plugs into it, so you're not paying twice).
It's a layer on top of what you already own, not a new meter. Plus the Boardroom has full token-optimisation tutorials.
Wrong: "Agentic AI is expensive and only for big companies."
Right: The whole reason it was expensive was cloud tokens per loop. On a fast free local model, the loop costs nothing — so a solo operator can run the same agentic workflows a big team would.
Wrong: "A small local model can't really do agent work."
Right: The runs above are real — Gemma 4 used its shell tool, read a file, critiqued its own draft. For triage, summarising, watching and drafting, a fast local model is plenty.
Wrong: "I'd have to be technical to set this up."
Right: It's a short config file and one command — hermes chat --profile gemma-speed. Inside the Agent OS it's already wired; you just pick the workflow.
158 pages of members already automating with agents — real businesses, real wins, in their own words.
Read the 158-page wins doc →A fast, free, local model is the missing piece that makes agentic workflows actually affordable to run all day. That's the difference between AI as a novelty and AI as a worker.
Inside the AI Profit Boardroom you get the full Agent OS — the fast local engine, these workflows pre-wired, Agent Kanban for overnight crews, cron watchdogs, memory that knows your business, every CLI you already pay for in one dashboard, a 30-day roadmap, daily tutorials, coaching calls, and 3,900+ founders across 38 countries building alongside you. Every new update gets folded in the week it lands.
It's the operating system I run a seven-figure business on. You get the whole thing.
Get the Agent OS →Set up the gemma-speed profile and run one loop tonight. I'll see you in the next one.