The stop button for AI agents just got cheap

The short answer: checking every action an AI agent takes used to be too slow and too expensive to do all the time. This week it stopped being either — and that changes how you should put agents into production.
TL;DR
- What changed: OpenAI announced a Decisions API that picks from fixed options instead of writing text. One reported demo put agent monitoring at $2.94 versus $372 with a frontier model.
- Why it matters: a check before every agent action is now affordable. Let the cheap check route the risky ones to a person.
- Our numbers: we use Jev, a decision model, in our own AI coding tools at about $0.0001 per check. After builds the agents called “done”, it found missing requirements in 12 of 17; only 1 passed first time.
- The limit: it routes, it doesn’t judge. Our reviewer overruled the model panel about 1 time in 4.
- Do this week: list the actions your agents take that customers would notice, and put a proceed / block / ask-a-person check in front of each.
What this means is our weekly note on the AI news that matters to teams shipping AI. This week: one idea, three headlines, and our own numbers.
What changed this week?
On September 30, TechCrunch reported that OpenAI announced a Decisions API at its Dev Day. Instead of writing an answer, the model is given a fixed set of options — categories for an image, or allowed agent behaviours — and picks one. OpenAI’s pitch: focusing the model on the choice makes it “extremely fast” while keeping its safety protections.
The same article covers Jev, the decision model TypeSafe AI released on September 18 that the API resembles, and a hackathon demo in which monitoring agent actions cost $2.94 with Jev versus $372 with a frontier model. That is one demo, not a benchmark, but the direction is hard to miss. TechCrunch ties the timing to incidents where OpenAI’s own agents misbehaved on the open internet.
The same day, NVIDIA and CoreWeave announced tooling for the other end of the problem: CoreWeave Agent Lens, which they say cuts the time to detect agent failures by 20% and the cost of fixing them by 50%.
Why does a cheap check matter so much?
Because the safety question for agents is not “is the model smart?” It is “who looked before it acted?”
Until now, most teams had two bad options:
| Approach | Problem |
|---|---|
| Ask a frontier model to review every action | Slow and expensive, so teams sample or skip it |
| Let the agent act and review the logs later | You find out after the email went out or the money moved |
A decision model that answers in under a second for cents makes a third option normal: check every action, and only escalate the uncertain ones to a person. That is the pattern that separates agents that demo well from agents that survive production — the gap we wrote about in Fast to Build, Slow to Ship.
How do we use Jev in our own tooling?
We have run Jev since the week it launched, inside the skills our AI coding agents follow on every Viaknox project. Jev never writes code or prose. It only answers narrow questions — yes or no, a score, one option from a list — with a confidence. Plain code then decides what those answers mean, and every rule fails closed: if Jev errors or is unsure, the work is escalated, never waved through.
| Where Jev sits | The question it answers | What happens next |
|---|---|---|
| Frontier gate — before an agent presents options or a “fix it later” plan | Is this worth a second opinion? Is the recommended option a shortcut? Is the brief missing context? How high are the stakes? | Sends the decision to a panel of four frontier models and sets how hard they think. After the build, Jev checks the change against the panel’s checklist. |
| PR review loop — before and during code review | How risky is this change: normal, sensitive or critical? For each reviewer-bot comment: fix it, reject it, or escalate? | Rules can only raise a route, never lower it. Anything touching auth, secrets, data loss or payments goes to a person. Jev never approves a deploy or a destructive command. |
| Session handshake — when one AI session hands work to the next | Can a fresh session act on this handoff alone? Is every decision recorded with its alternatives and reasons? Are there secrets in it? | The handoff is rewritten until it passes. A regex secrets scan runs as well, whatever Jev says. |
| Adversarial review — our open-source release gate | Not used, on purpose. | The final ship or no-ship verdict is computed by a script from recorded test and review results. We keep every model, Jev included, out of that call. |
From our logs, September 18–30, 2026:
| What we track | Result |
|---|---|
| Jev triage checks logged | 30, averaging $0.0001 each |
| How hard it told the panel to think | High 12 · Medium 17 · Low 1 |
| Decisions reviewed by the four-model panel | 53, costing $29.21 in total (about $0.55 each) |
| Person agreed with the panel’s pick | 73% |
| Person agreed with the coding agent’s own pick | 64% |
| Finished builds Jev checked against the panel’s checklist | 17 |
| …that passed first time | 1 |
| …with at least one agreed requirement missing | 12 (about 5 missing on average) |
| Health check on Sep 30 | Answered in 715 ms |
What value do we actually get from Jev?
The biggest win came after the build, not before it. Our coding agents reported “done” on builds that Jev then found were missing, on average, five of the requirements the panel had agreed. Only one build in 17 passed first time. Without that check, those gaps would have reached review — or production.
It rarely lets us skip a review — and that is fine. The decisions that reach our gate are already non-trivial, so Jev escalated nearly all of them. Its real job there is choosing how hard the panel should think, which is what controls cost and waiting time.
It is cheap enough to run everywhere. At about a hundredth of a cent per check, against roughly 55 cents for a panel review, we can put a check at every hand-off — between sessions, on every bot comment, after every build — not just on the big decisions.
It routes; it does not judge. An independent panel beat the agent that proposed the change at matching the human’s call, by nine points — agents grading their own homework is still a bad idea. And the person overruled the panel about a quarter of the time, usually on trade-offs only a product owner can weigh. Jev decides what reaches a person. It never replaces the person.
Caveat: these are field notes from our own projects over two weeks — a small sample, not a benchmark.
What else happened?
- Small businesses are already using agents. OpenAI reported that about 4 million employees at companies with fewer than 500 people used its products in one September week. It also announced a partnership with America’s SBDC to train about 150 advisors and reach at least 1,000 small businesses through workshops. Small teams are exactly the ones without a security department — so they benefit most from cheap, built-in checks.
- Attackers target the models, too. OpenAI described disrupting a coordinated campaign, starting in July, that used more than 15,000 accounts to extract protected model reasoning. Its own conclusion: this “requires layered controls and continual adaptation.” The same is true of your agents.
- Own or rent? A sponsored piece from HPE in MIT Technology Review argues that once AI workloads are steady and predictable, owning capacity beats paying per request. Note the sponsor — and its own caveat that ownership “only makes sense when an enterprise can keep capacity productive.” For most teams we work with, usage is not steady yet.
What should you do this week?
- List every action your agents can take that a customer would notice — emails, payments, record changes, public posts.
- Put a decision check in front of each one with three options: proceed, block, or ask a person. Log every answer.
- Review the “ask a person” cases weekly. They tell you where your agent’s judgment and yours differ — which is where the next failure will come from.
If you want help designing that gate for your own agents, that is the Architect stage of our consulting work: guardrails designed before the first model call, not after the first incident.
The full, live list of this week’s AI headlines is on our News page.
Questions about this article
What is a decision model?
A model that is only allowed to pick from a short list of options you give it, such as allow, block or ask a person. Because it never writes free text, it can answer in well under a second and costs a fraction of a general-purpose model. TechCrunch reported one monitoring job costing $2.94 this way versus $372 with a frontier model.
Does a cheap check replace human review?
No. In our own gate, a person makes the final call on every escalated decision, and they disagreed with the model panel about a quarter of the time. The cheap check decides what reaches a person; it does not decide instead of them.
We are a small business. Is this relevant yet?
Yes, if you are letting an agent send email, move money, change records or post publicly. OpenAI says about 4 million employees at companies under 500 people used its products in one September week. Once an agent can act, a pause-or-proceed check before each action is the cheapest insurance you can buy.
Where does Viaknox use Jev?
In the skills our AI coding agents follow — a gate before options are presented, a verification pass after each build, code-review triage, and session handoffs. We deliberately keep it out of the final release verdict, which our open-source adversarial-review gate computes from recorded test results.
What is "What this means"?
A short weekly note from Viaknox on one or two AI headlines — what changed, what it means for teams putting AI into production, and our own numbers where we have them. The full list of this week's headlines is on our News page.
Sources
- TechCrunch — OpenAI's Jev clone could help the frontier lab stop its swarming agents (Sep 30, 2026)
- NVIDIA — From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI (Sep 30, 2026)
- OpenAI — Helping small businesses put AI to work (Sep 29, 2026)
- OpenAI — Disrupting a coordinated model-distillation campaign (Sep 30, 2026)
- MIT Technology Review (sponsored by HPE) — Making AI an asset, not an expense (Sep 29, 2026)
Have a decision or a system to get right?
Consulting: Assess, Architect, Deliver — with your team, not around it.