Why agent workflows need a graph, not a queue
What broke in our first production agents, why a queue could not fix it, and how a graph of typed steps gave us retries, idempotency and an audit trail.

If you’ve built with AI long enough, you’ve seen this moment: the agent opens a pull request, writes its own tests, stamps “Done”… and something obvious is still wrong.
In our demo — the one you can check in the public repo — the story was dead simple: orders of 10 or more items get 10% off. The AI wrote the code, wrote the tests, and confidently declared victory. And an order of exactly 10 items still paid full price.
The fix wasn’t “use a bigger model.” It was a rule: a story isn’t done until the AI shows evidence that each acceptance criterion works — and that evidence must be checked by something other than the AI that wrote the code.
That’s what story-gate does! It’s free, open source, and already in our build stack.
This wasn’t hypothetical. It was a staged demo in a public repo.
The story: orders of 10 or more items get 10% off.
The agent wrote: if qty > 10:
It tested 9 items and 12 items. Both passed. It marked the story done.
The bug was a single character. > instead of >=. At exactly 10 items — four dollars each — the customer paid 40.00 instead of 36.00. Every check the AI ran on its own work said “looks good.”
story-gate reran the story’s scenarios on GitHub, against the real code. Pull request #5 came back with: “DONE gate FAIL… exactly 10 items get the discount (output did not match).”
The independent judge scored “implementation matches acceptance criteria” at 0.04 out of 1. Merge blocked.
The agent fixed it in pull request #6: changed one character, added a test for exactly 10 items, and wrote down the lesson — test the exact boundary the story names. Every scenario passed. A human approved it.
| Pull request #5 | Pull request #6 | |
|---|---|---|
| The AI says | “Done” | “Done” |
| Its own unit tests | Pass | Pass |
| story-gate runs “exactly 10 items” | Wrong total | Correct total |
| Merge | Blocked | Allowed after a person approves |
Two PRs. Both said “Done.” Only one actually was!
We learned this the slow way — building our own products with AI.
Spec kits like GitHub’s Spec Kit, OpenSpec, and BMAD are great for planning. They help the AI think before it builds. But they guide; they don’t enforce. Nothing stops a merge when the code doesn’t match the plan.
Our first real failure wasn’t bad code — it was thin stories. Acceptance criteria like “discount works” or “user can save settings” sound fine… and are completely untestable. The AI can always claim it met a vague goal. That’s how slop gets through — features that demo well but break at the edges.
The Production Readiness Kit checks the whole product before launch. But by launch day, hundreds of small stories have already merged. If each one was “done” on the AI’s word, the readiness check becomes a late-stage fire drill.
Governance must run inside the workflow — not at the finish line. story-gate applies that idea at the smallest scale: every story, every day.
Proof means evidence against the acceptance criteria. Not vibes. Not confidence. Not “the tests passed.” Actual evidence: a passing test and a recorded run for each goal.
Proof doesn’t mean bug-free code. story-gate only checks what the story asks for. If the story is wrong or missing something, the gate will happily prove the wrong thing.
That’s the point. It forces the hard question up front — what exactly must be true when this is done? If you can’t write that down, the story isn’t ready.
story-gate puts four gates on each story:
Three design choices matter:
story-gate isn’t designed to replace your tools. It sits beside them. Keep your spec kit. story-gate reads its scenarios and checks the code against them.
If you’re a founder or CTO building with Claude Code, Cursor, Codex, Gemini CLI, Hermes, or VS Code agents — and you vibe-code your way through features — you know the pain!
The AI is fast. The AI is confident. The hard part is knowing whether “done” is actually done!
story-gate is built for you. It forces clarity before the build, shows real evidence after the build, and keeps the AI honest. You don’t need to read every line of code. You read a one-page report and click Approve.
Skip it for weekend hacks. Use it for anything that matters. The checks add a few minutes per story, plus one small judge call per check through OpenRouter.
We publish what we use. story-gate came out of our own builds, and we’re putting it behind every Viaknox product.
It’s a pilot. v0.9 is signed and tested on Linux, macOS, and Windows, but commands may change before 1.0. The demo in this post is staged — not a year of production data.
Our bar is simple: if a story can’t show evidence for its acceptance criteria, it isn’t done. That applies to our code too.
You can watch story-gate catch a bug without touching your own projects. story-gate try builds a throwaway example in a temp folder, with one story and one real bug, and opens the validation page. No GitHub. No judge key. No network.
Mac or Linux:
curl -LsSf https://raw.githubusercontent.com/SathiaAI/story-gate/v0.9.0/install.sh -o /tmp/story-gate-install.sh && sh /tmp/story-gate-install.sh
Then open a new terminal (so it finds story-gate) and run:
story-gate try
Windows (PowerShell):
powershell -ExecutionPolicy ByPass -c "irm https://raw.githubusercontent.com/SathiaAI/story-gate/v0.9.0/install.ps1 | iex"
story-gate try
Read the installer first if you want — it’s short.
Then tell us what broke. We want the bugs it missed, the checks that felt like noise, and the setup steps that made you swear. → Join the story-gate discussion on GitHub
No. They work at different levels. story-gate keeps each change honest, story by story. The kit checks the whole product once before launch — security, backups, monitoring, legal and the rest.
No. Proof means evidence against the acceptance criteria the story states. If a criterion is missing, story-gate can't catch the gap — which is why the READY check makes you write testable goals before any code.
No. You need a GitHub account and an AI coding tool. The AI writes the stories and runs the checks. You read a one-page report and approve on GitHub.
Claude Code, Codex, Cursor, VS Code (Copilot agent), Gemini CLI and Hermes get live checks; Windsurf is partly verified; cloud agents get the pull request checks. It is MIT-licensed and free — the only cost is the judge calls.
Yes — to the judge you choose (OpenRouter by default), with that story's evidence and code changes. You can run without a judge in "objective mode", which still runs your tests and scenarios and still needs your approval.
Articles and announcements, delivered by Substack.