Released Sep 15, 2026 · TypeSafe AI · now in beta on OpenRouter

Jev AI Just Changed AI Agents Forever

It doesn't write. It decides — and tells you how sure it is.

Jev AI just changed AI agents forever.

It was built by one of the people behind the research that became ChatGPT.

It makes your AI agents decide up to two hundred times faster.

It also tells you how sure it is, so you know exactly when to trust it.

People have already used it to sort inboxes, score leads and rebuild a whole website's links in seconds.

Stick with me, because I ran it on my own website and I'll show you what happened.

A brand-new kind of AI model: it doesn't write, it decidesThe situationwhat's happeningright nowA questionwith set answersJev decidesin milliseconds+ how sure it isa number youcan act onThe writing brain keeps writing. The red lights get their own brain.
The actual sources ↓
§1The person behind it

Built by one of the people behind the research that became ChatGPT

Diogo Almeida's last four yearsOpenAIhelped build theresearch behind ChatGPT2 years quieta new way totrain modelsSep 15, 2026TypeSafe AIreleases JevThen this week, he released Jev
"At OpenAI, I helped build the methods that made language models useful at following instructions and talking with people. That work ended up as the research behind ChatGPT."— Diogo Almeida, founder, TypeSafe AI · launch essay · Sep 15, 2026
§2The numbers he put on it

Up to 200 times faster. Up to 400 times cheaper.

End-to-end response timeA normal frontier model3 – 329 secondsJev70 – 500 msFast enough to sit inside your software
What a million tokens costsNormal models · input$0.20 – $10 / M tokensJev · input$0.042 / M tokensJev · outputfree193.6× faster and 444.6× cheaper on TypeSafe's own workflow tests
numbers: TypeSafe AI launch essay + docs
§3Then somebody plugged it into Claude Code

A Claude session went from nearly one million tokens to 86,000

One second. Same session. A fraction of the context.BEFORE · nearly 1,000,000 tokensAFTER · 86,000 tokensPeople reported context going from 90% full to 9%
§4So what actually is this thing?

Why is everyone building with it inside a week?

Tweet · the reaction

“A genuine phenomenon”

Obie Fernandez wrote the book on putting AI into real software. He says ordinary programmers are learning to plug fast, cheap intelligence into normal code, and he expects every big lab to ship a Jev-style model within months. Shopify's founder replied “totally agreed”.

§5What Jev is not

It is not a chatbot. It does not write. At all.

What you can and can't ask it✕ Write an emailcan't✕ Write codecan't✕ Explain its pickcan't✓ Make decisionsfast ones✓ Thousands of themin parallel✓ Say how sure it isevery timeIt doesn't write. It decides.
THINKING IT? "If it can't write anything, what use is it to me?"

Most of what slows an agent down is not the writing. It's the pile of tiny decisions between the writing.

That pile is the part Jev takes over.

§6How you use it

You hand it the situation and a question. It picks — and tells you how sure it is.

Situation in. Decision out.The situationwhat's happeningright nowA questionwith setanswersJev picks onein milliseconds+ how surea real numberyou can act onThink multiple choice, not essay writing
§7System One

The fast part of your brain: you see a red light, you just stop

Same red light. Two very different brains.System Oneyou just stopa fraction of a secondSystem Twowrites a paragraphabout the red lightYour agents use the paragraph brain for the red lights too
§8The problem

A big model, a full history, seconds of thinking — for a one-word answer

Every tiny question goes to the big modelDo I have enough sources yet?Write now, or keep looking?Good enough to send?Which tool do I use next?A big modelreads everythingwrites its reasoningSeconds go by. Tokens get spent. The answer was one word.
That's the gap Jev fillsOLD WAYseconds eachEvery tiny question → big modelReads the whole history againWrites out its reasoningYou pay for every word outOne question at a timeNEW WAYmilliseconds eachTiny questions → JevReads the situation onceNo writing, just the pickOutput is freeA whole batch at onceOnce you see the gap, you see it everywhere in your setup
Seconds go by. Tokens get spent. And the actual answer was one word.
§9Three kinds of question

Choice, Score and Noul cover almost everything

real Jev run · my machine · replayed at reading pace

What you're watching: one real lead email goes in with three questions — a Choice, a Score and a Noul — and all three answers come back from a single request, with the real probabilities, latency and cost.

§101 · Choice

Pick one from a list — with a probability for every option

“Which tool should the agent use next?”web_search98%read_docs1%write_file1%none0%The pick, every probability, and one overall confidence: 0.98
real answer from my research-agent test
§112 · Score

Rate it against levels you define — and it can land between them

“How well will this outreach message land?”0 · ignoredwill annoy them1 · genericmight get a glance2 · relevantdecent chance3 · specificstrong chanceReal answer from my lead test: 2.4 out of 3 — between two levels
§123 · Noul

A yes or no — returned as the probability of yes

Three real Noul answers from my tests0.99“sounds urgent?”0.47“needs a human?”0.05“is `ls` risky?”Near 0.5 means it genuinely doesn't know — and that's useful
§13The part most people skip

It tells you how sure it is — so you can set a line

My 60-email test: every dot is one real decisionmy line · 0.6016 below the line → come and ask me44 above → carry on0.0 unsurecertain 1.0All 6 wrong picks (red) sat below the line
Tweet · a line in the wild

“75%+ → show me. Everything else gets ignored.”

Knowix built a content radar. Every hour it pulls posts and news, then Jev judges each topic against five questions. Jev writes nothing. Anything it's 75% sure about lands on a shortlist, and the rest is logged and ignored. That's the set-a-line idea running in real life.

§14Why that changes everything

The problem was never the AI being wrong. It's not knowing when it's wrong.

Same 60-email test, read the other way54 / 60right first time6 / 6mistakes under the line0wrong ones sent aloneThat single idea changes how much you can safely hand over
It's not the AI being wrong. It's not knowing when it's wrong.
§15Lots of questions at once

Five questions instead of one barely changes the speed or the cost

real Jev run · my machine · replayed at reading pace

What you're watching: one research situation, five tiny questions in a single request — all five answers come back together. Then my timing test: one question took 377 ms, five questions took 322 ms (median of six runs each).

§16What people built in the first few days

The results are specific, and they're all public

Tweet · the playful one

50 games at once, for less than a cent

Max Blade pointed Jev at Subway Surfers. It plays at superhuman speed and runs fifty games at the same time. His own caveat is the important part: Jev does not replace the big models, it opens up a new kind of job.

§17Build 1 · research sorting

1,018 research papers sorted for eight cents

Hassan's paper sorter1,018AI research papers24possible topics256 msmiddle time per paper$0.08total costSummarise each paper, then let Jev pick the category
How he set it upPapertitle + summary24 topicsas the optionsJev picksthe categoryA quarter of a second per paper, start to finish
§18Build 2 · email sorting

500 emails sorted in seconds for three and a half cents

Riley Brown's inbox sorterreplyresearch this firstwaitflag for meOne emailgoes in asthe situationEvery answer points the email at a different folder or helper
real Jev run · my machine · replayed at reading pace

What you're watching: I ran the same idea on a 60-email test inbox I wrote. 2.53s wall clock, $0.0010 total. The small number on each email is Jev's confidence.

§19Build 3 · lead scoring

700 leads scored in 40 seconds — and the mismatches flagged

Romàn's outreach check700leads + their messages40 sto score them all$0.09total costIt points at the messages aimed at the wrong person
real Jev run · my machine · replayed at reading pace

What you're watching: my 40-lead test list with 8 deliberately mismatched messages hidden inside. Jev flagged 8 as “wrong person” — and all 8 were the planted ones. 1.52s, $0.0007.

§20Build 4 · browsers

Browser Use found flights in seven seconds, for under half a cent

Jev decides, the small model types, the browser clicksPage changesafter a clickRebuild listwhat's clickableright nowJev picksfrom thefresh listSmall modeltypes the textwhen neededBrowserclicks, thenrepeatThe speed came from cleaning up the loop, not a bigger brain
They measured it properly — same models on both sidesBrowser commands · before1,092Browser commands · tuned101Task time · before9.45 sTask time · tuned7.09 s (−25%)Middle of three matched runs each
honest note: it finds flights, it doesn't book themthe 7-second clock starts after the first page loadtheir code checks the result separately after Jev says “done”
Tweet · the step-by-step

“Give those decisions their own model”

Codila's long article walks the whole setup: find the part of your agent Jev can take over, try one decision in the Playground, then run Browser Use's agent — the one that found flights in 7 seconds for $0.0039. The clearest line in it: your writing agent keeps writing, and the small judgments between become a separate part you can inspect and price.

§21Build 5 · the one closest to what I run every day

586 pages re-linked in 45 seconds. Claude Opus 5 got through 21.

Borja's internal-link rebuild — same 586 pages, same 45.1 secondsJev · pages done586Claude Opus 5 · same clock21584 links placed · 139 pages left alone · 21 cents
real Jev run · real agentos.guide pages

What you're watching: I ran it on this website. 257 real agentos.guide pages, 1,542 yes/no decisions, 5.8 seconds, $0.011. It placed 356 links and left 92 pages alone because nothing honestly fit.

THINKING IT? "Internal linking tools always force a link onto every page."

That's exactly what goes wrong. “Nothing fits here” is a scoring job, not a writing job.

On my own site Jev left 92 pages alone. A tool told to “add links” would have stuffed all 257.

“Nothing fits here” is a scoring job, not a writing job.
§22Why this one matters to me

Hundreds of small yes-or-no decisions, every single week

Content across several sites at onceSite 1 · link? yes / noSite 2 · link? yes / noSite 3 · link? yes / noSite 4 · link? yes / noSite 5 · link? yes / noNew articlewhich old pagesshould it point at?Always the slow, expensive part — until now
§23 Quick word

Use Jev inside your business — without touching a line of code

We've built an agent operating system you can download as a zip file. The exact job Jev does is the layer we walk through, step by step, on the coaching calls.

The Agent OS zip file — plug in your Claude, your Hermes, your OpenClaw, all of it, into one dashboard with one shared memory
The decision layer, walked through live — routing work between your agents, scoring leads, sorting your inbox, deciding what needs your eyes
Four coaching calls every week — bring your own setup and ask
Daily tutorials — new step-by-step videos as tools like Jev land
A 30-day roadmap — your first agent running and pointed at bringing in customers
3,900+ business owners already inside — a lot of them had never used AI before they joined
Join the AI Profit Boardroom →Inside the AI Profit Boardroom · skool.com/ai-profit-lab
4 live calls a week · daily tutorials · used in 38 countries
THINKING IT? "Doesn't running an agent operating system burn a fortune in tokens?"

No. The everyday work runs on free local models on your own machine, free APIs slot in for more, and the heavy work drives the CLIs you already pay for — your Claude subscription already includes the Claude CLI, and the Agent OS plugs straight into it.

A decision model like Jev pushes that even lower, and the Boardroom has full token-efficiency tutorials.

§24Feature 1 · routing between models

“Choose the cheapest model that can finish this job”

You describe each model in plain English. Jev routes.Cheap + fast · lookups, small changesMid · longer writing, routine workExpensive + smart · hard decisionsEvery requestJev reads itand picksThe probabilities stay on file, so you can check the routing later
real Jev run · my machine · replayed at reading pace

What you're watching: six real requests, one plain-English instruction. The timezone question and the rename go to the small model, and the strategy plan and the race-condition hunt go to the frontier model.

LangChain shipped this as an official Jev integration
§25Feature 2 · safety checks on tool use

The guard rail that used to be locked inside someone else's product is now yours

A wrapper you put around your own agentYour agentwants to runa tool callJev checks ithow risky is this?Allow · ask · blockbefore it runsIt watches every tool call and blocks the risky ones first
real Jev run · my machine · replayed at reading pace

What you're watching: eight tool calls an agent might try while building a report. ls and git status pass, emailing all 4,000 members stops to ask me first, and rm -rf, DROP TABLE, a force-push and a piped install script get blocked.

§26The context one

Score every tool call in the history. Drop the ones that don't matter anymore.

What people reported~1M → 86Ktokens, in about a second90% → 9%context full30–60%cuts were commonThe loudest reaction of the week
real Jev run · a small test history I built

What you're watching: my own small version. A 14-step agent history, one Jev request scoring every step against the current goal. 132,000 tokens became 10,900 in 362 ms.

§27The honest pushback

Theo: cleaning up history is not the same as filtering it

Two different jobsSCORE + DROPa filterScores tool calls one by oneThrows away the low scorersCan lose the reasoning trailLong tasks can break laterCOMPACTIONa recordKeeps a proper recordOf what happened, and whyThe trail stays intactBuilt for long tasksHe's right about the risk
§28Where I land

Right about the risk. And the direction is correct.

Your agent's history is full of stuff that mattered for two minutesUse them to delete thingsOr to decide what to summarise firstRelevance scoresin one secondThat's the open question — and it's a week old
§29What Jev can't do · 1

It can't write anything. Jev only picks.

Three jobs, three ownersA normal modelresearchesand writesJevroutes, scores,approves, escalatesYour codesaves the file,does the thingThis is where people are going to trip up
§30What Jev can't do · 2

Confident is not the same as correct

Browser Use does exactly thisJev says “done”0.97 sureSure ≠ rightand it provesnothing happenedYour code checksdoes the fileactually exist?✓ Verifiednow carry onAfter Jev says done, they check the outcome separately
§31What Jev can't do · 3

It can be tricked

real Jev run · my machine · replayed at reading pace

What you're watching: I tested it. A clean refund complaint gets flagged for the owner with certainty. Paste a hostile note into the same email and “archive it” jumps from 0% to 25%. One hardening line in the question pulls it back to 97%.

§32A detail that's easy to miss

Jev doesn't see the name you give your question

real Jev run · my machine · replayed at reading pace

What you're watching: the same draft, containing a password and a client's contract value, asked three ways. A tidy field name with an empty question gets a shrug (0.41). A real question catches it (0.99). The same real question under a junk name gives the identical answer.

§33Same goes for what you feed it

The quality of the decision follows the quality of the situation

real Jev run · my machine · replayed at reading pace

What you're watching: same question, three situations. Tell it only “the researcher finished” and it says “write the draft” at 87% — a guess. List the sources and what's missing, and it says “keep researching” at 99%. Close the gaps and it says “write” again, this time for a reason.

§34Cost

You pay for what you send in. What comes back is free.

The pricing shape is unusual1,000 tokensof contextper decision× 10,000decisions= 10M tokens inat about 4¢per million≈ 42 centsoutput: freeIn beta on OpenRouter — the same place you reach everything else
My own receipt$0.013everything you saw me run today1,787real decisions in those runsEmails, leads, the whole-site link map and the history test
§35The number to actually watch

Cost per finished task, not cost per decision

Cheap routing that routes badly is expensiveCHEAP + WRONGexpensiveA fraction of a cent to decideSends the agent down the wrong pathThe big model does the wrong workYou pay for all of it, twiceCHEAP + RIGHTcheapA fraction of a cent to decideRight path, first timeThe big model only does real workTask finished, onceMeasure the finished task
§36What's changed · one

Writing and deciding used to be the same job. Now they're separate.

Same model, same price — not anymoreThe big modelresearches, plansand writesJevroutes, scores,approves, escalatesYour codedoes the thingHere's the shape of what's actually changed, in three parts
§37What's changed · two

Decisions used to queue up. Now a whole batch comes back together.

Ask them all at onceWhich step next?How urgent is this?Does this need approval?Which tool?Send, or review?One situationone requestThe waiting in the middle of your agent mostly disappears
§38What's changed · three

Confidence became a number you can act on

Set a line. Above it, run. Below it, come and ask.Set a linesay 0.80Above itrunBelow itcome and askThat's what makes it safe to leave running while you go to the gym
§39Why this matters more to a business owner

The thing holding people back is trust, not capability

It can write the email. Will it send the wrong one?62%“I'm 62% sure — you should look at this one”That solves a trust problem, not a technical one
“I'm 62 percent sure — you should look at this one.”
§40The honest state of things

Days old. Marked experimental. Smart people are still arguing.

The safest use right now: the boring, repeated decisionSortingthe safe zoneScoringthe safe zoneRoutingthe safe zoneFlaggingthe safe zoneWhich is usually where the real gains are hiding anyway
§41“This sounds like developer territory”

Look at what these jobs actually were

Those are admin tasksSorting 500 emailsadminScoring 700 leadsadminWhich page links whereadminWhich helper gets the jobadminYou're doing them by hand, paying someone, or skipping them
Wrong: “This is developer territory. It's not for a business like mine.”
Right: Every job on this page was an admin task — sorting, scoring, linking, handing out work.
Wrong: “I don't have time to set any of this up.”
Right: The tools got cheap and fast enough this week that “no time” stopped being the reason.
Wrong: “I'd need to be technical to use it.”
Right: Every result here came from someone describing a decision in plain English and giving the model a list of options. That's the whole skill.
Don't take my word for it

Members post their wins every day — agency owners, ecom founders, course creators, solo operators across 38 countries.

Read the 158-page wins doc →
§42 Your move

Build the decision layer into your business — with help

The people who separate the writing from the deciding are going to run faster agents for a fraction of the cost.

The Agent OS as a zip file — Claude, Hermes and OpenClaw in one dashboard sharing one memory, updated as things like Jev land
Four coaching calls every week — ask what should be routed, what should be scored, and what should still come to you
Daily tutorials — wiring a decision layer into your lead sorting, your inbox and your content
A 30-day roadmap — so you know what to build first
A prompt library — every prompt, ready to paste
A member map — find people near you doing the same thing
3,900+ business owners inside — plenty started with zero AI experience
Join the AI Profit Boardroom →Inside the AI Profit Boardroom · skool.com/ai-profit-lab
Start with one decision you make over and over · 38 countries · 4 live calls a week
§43Start here

Start with one decision you make over and over

The whole methodOne decisionyou make overand overA list of optionsin plain EnglishMeasurewhat happensTake the next oneand repeatEveryone else keeps paying frontier prices for a one-word answer