Ox Alpha, the mystery AI model, just dropped — and nobody knows who built it.
A frontier model appeared out of nowhere two days ago, wearing a mask.
It holds a million tokens at once, it can watch video, and right now it costs nothing.
Independent testers ran it against Claude and GPT, and the scores turned heads.
Then internet detectives found fingerprints hidden inside the model itself.
By the end of this page you'll know the whole story before the mask comes off.
These are real recordings from the Free AI Coder section of my Agent OS — Ox Alpha wired in as an engine, the same day it appeared.
What you're watching: the Free AI Coder engine dropdown switching from OmniRoute to Ox Alpha · stealth 1M — the status chip flips to "Gateway live" the moment OpenRouter confirms the model is up.
What you're watching: I ask it straight — "Who built you?" It answers: "I was built by an undisclosed organization, and I'm the model known as ox-alpha." Even the model keeps the mask on.
What you're watching: one build prompt, and Ox Alpha writes a full animated page that renders live in the preview pane — then gets saved to my workspace. It even added the caption "IDENTITY WITHHELD" on its own.
On August 20th a listing called Ox Alpha showed up on OpenRouter and OpenCode — no company, no logo, no announcement, just the word "stealth" as the provider.
This is the post that started it. OpenCode lists the whole offer: 1M context, multi-modal, zero data retention, "generous rate limits, near unlimited usage" — and the line everyone quoted: "We have capacity for 100T tokens per day, lets see what you can do." It's sitting at 5.8 million views.
These four facts don't normally go together.
Connect once, and hundreds of models pull up to the same platform — that's how a masked model can appear overnight and reach everyone.
The listing calls it a reasoning model for coding, sustained agentic work, and production workloads — plain English: an agent that works for hours without losing the plot.
Here's the official wording from OpenRouter: "a frontier model built for efficient coding, sustained agentic work, and real-world production use" — with the 1M token window and text, image and video input spelled out. Their follow-up note adds the part that matters: it's free, and this provider does not train on your prompts.
The context window is the model's desk — a million tokens means months of notes, reports and plans on that desk at once, all visible while it works.
Text, images, and video in — the video part is rare, and it becomes the biggest clue in the detective story later, so remember it.
Very few companies on Earth have that kind of computing power sitting around, and that single number cuts the list of suspects way down.
Researcher Ben Davis ran it through ten real software engineering tasks on DeepSWE — here's how the run landed.
KC posted the comparison that went around: "gpt-5.6-sol → 52% · Fable → 65% · Ox Alpha → 80%+. People are starting to get very confused." It quotes Ben Davis's original run — his own words: "I am very confused." Read those names again: Claude, GPT — beaten by a model with no name on it.
There are no official benchmarks and these community numbers aren't independently verified — treat them as a strong early signal, not proof.
Nous Research plugged Ox Alpha into their Hermes Agent project, and the code editor Zed integrated it — on its first day online. That tells you something benchmarks can't.
This matters to me personally, because Hermes is one of the main agents inside my own Agent OS. Nous didn't just test the anonymous model — they put it in their portal and wrote: "We have capacity for 1 quadrillion tokens per day. Let the tokens flow." When a team whose tools I run daily routes real work through a masked model, I pay attention.
I plugged Ox Alpha into my Agent OS the same day through OpenRouter, right alongside Hermes and Claude — and here's the page it built for me, running full screen.
What you're watching: the page Ox Alpha built inside my Free AI Coder, opened full screen — twinkling starfield, pulsing golden orb, and its own "IDENTITY WITHHELD" caption. The mystery model has a sense of humour.
No — that's the biggest myth about it. This whole guide runs on a free model, the everyday 90% runs on free local models and free APIs, and for frontier work it drives the CLIs you already pay for — your Claude subscription already includes the Claude CLI.
It's a layer on top of what you already own, not a new meter. And inside the AI Profit Boardroom there are full token-optimisation tutorials, so usage never becomes a worry.
You just watched Ox Alpha running inside mine. This week specifically, we're testing Ox Alpha inside the Agent OS — so when a stealth model drops, you test it in your own system the same day instead of watching from the sidelines.
Nobody claimed Ox Alpha, so the community went hunting for the tiny habits a maker can't easily hide — and Ben Davis found three of them.
Davis fed Ox Alpha four controlled videos and counted the tokens — the counts matched Zhipu's GLM-5V-Turbo token for token, while every other candidate had a clearly different signature.
Across 25 prompts, Ox Alpha's token counts matched GLM-5.3 exactly apart from a constant 75-token hidden wrapper — matching that precisely requires an identical vocabulary.
Ox Alpha rejects audio input with the same behavior as GLM-5V, while Xiaomi's MiMo — the leading rival theory — happily accepts audio. They even checked the emoji rate: it matched Zhipu's models.
Stack the clues and Davis puts it at 99% a Zhipu GLM 5 model, with other analysis around 90% that this is the next-generation multimodal model — possibly what people will call GLM 5.5. Officially, it is still unconfirmed.
The previous four were all eventually claimed by Chinese labs — anonymous debut, a burst of free traffic, then the company steps forward.
Launch masked: if it wins, take the mask off and enjoy the applause — if it stumbles, fix it quietly and nobody connects it to you.
Zhipu shipped GLM 5.3 on August 14th with no image or video input — the number one community ask was a vision version, and six days later an anonymous model with GLM's exact fingerprints shows up seeing both.
Testers measured how fast it generates text: within about 6 percent of GLM-5V-Turbo — the size class where 100 trillion free tokens a day is a real stress test, not a fantasy.
A year or two ago this capability was locked behind enterprise contracts — now it shows up anonymously on a public platform, and anyone with an OpenRouter account can point their agents at it.
I've been feeding mine long recordings of my own screen work and having it write up what I did — that used to take hours by hand.
Ziwen wired Ox Alpha into their coding setup and described exactly this shift: "we stopped picking the files we think matter and started handing it the whole repo… this week is for the jobs that are usually too big." No API key, no billing, a million tokens of context — until the window shuts.
The listing says prompts and completions are retained by the provider — just not used for training — and the provider is anonymous, so keep client data on models with a name.
Wrong: "A new model every week — I can't keep up, so why bother starting?"
Right: The speed is exactly why people who build a system win. My Agent OS stays the same, my workflows stay the same — only the engine underneath swaps out. Ox Alpha appeared on a Thursday and ran my tasks by that evening.
Wrong: "I'm not technical enough for something called a stealth reasoning model."
Right: You just followed a story about tokenizer fingerprints and understood every bit of it. Choosing a model on OpenRouter is picking from a dropdown menu — you watched me do it in the first video on this page.
Wrong: "I'll wait until there's one obvious winner."
Right: The dust is not going to settle — five stealth models in six months is the new normal. You need one working system and the habit of testing new engines when they appear. That habit takes an afternoon to build.
Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries. Real businesses, real numbers, in their own words.
Read the 158-page wins doc →The free window is expected to close around the 27th, the reveal usually lands after the preview ends, and Zhipu's GLM 5.3 weights unlock after a safety review about two weeks out — the timelines line up for a very interesting end of August.
This week specifically, the AI Profit Boardroom is testing Ox Alpha inside the Agent OS — the system where your Claude, your Hermes and your OpenClaw share one memory and work together. Readers watch the mystery from the sidelines. Operators run the mystery model on their own work the same day it drops.