AI strategy

Fast to Build, Slow to Ship

Viaknox · ·15 min read AI strategy
Hand-inked city scene: a brand-new high-speed train waits at a station platform with passengers ready to board, but the track ends a few metres past the platform at an orange buffer stop. Fast to build, slow to ship.

We are building a suite of products at the moment, and building them is the easy part. Claude, Devin, Codex and Hermes have each taken a stab at pieces of it, and the demos come fast. What has not come fast is production. Before a single paying customer, we have spent over $5,000 on tooling and software alone to get security, reliability and privacy to a standard we would put our name on — scanners, review tools, monitoring, the lot — and that is before counting the founder’s time, which is the expensive part. AI made the build cheap. It made the checking expensive. Add it up and the bill starts to look like what a software firm would have charged, and a software firm would at least have brought experience with it.

That is the whole argument of this article, so it is worth saying plainly up front: AI made software fast to build and no faster to ship. A working demo now takes days. Getting that demo safe, reliable, supported and used still takes months. Most never get there. Below are the five places we have watched AI-built products stall, the readiness framework we now use on our own work, and a free kit that runs it inside whatever coding agent you already have.

Why is AI making software faster to build but not faster to ship?

Because it speeds up the making, and only the making. Everything around the code — the review, the sign-off, the on-call, the integration with systems you do not control — runs at the same speed it always did, and sometimes slower, because there is now more code to check and less confidence in who wrote it.

The published evidence points the same way, though not all of it is equally good.

Six numbers behind the build-speed paradox: −7.2% delivery stability, 19% slower yet felt 20% faster, 45% of AI code with security flaws, 46% distrust AI accuracy, 74% no tangible value, 21% have redesigned workflows.

DORA’s 2024 report models a 7.2% drop in delivery stability for every 25% increase in AI adoption (DORA 2024). It is a modelled estimate across many organisations, not a measured drop on any one team, so treat it as a direction rather than a forecast. The 2025 edition puts the same thing in words: AI “magnifies the strengths of high-performing organizations and the dysfunctions of struggling ones” (DORA 2025). That matches what we see. It does not fix a weak process; it runs a weak process faster.

The METR study is the one people quote most and trust least. Sixteen experienced open-source developers, working on their own repositories, took 19% longer with AI tools while believing they had been 20% faster (METR, Jul 2025). Sixteen is a small number and these were people who knew their codebases intimately, which is close to the worst case for an assistant. But the gap between felt speed and measured speed is the part worth keeping. We have felt it ourselves.

Then the quality numbers. Veracode found security flaws in 45% of AI-generated code samples across more than a hundred models (Veracode 2025) — a vendor test suite, not production code, but a vendor that sells scanning has no reason to understate it. Stack Overflow’s 2025 survey has 46% of developers distrusting AI accuracy against 33% who trust it (Stack Overflow 2025).

And the business numbers, which are the ones that should worry a founder. BCG found 74% of companies had yet to show tangible value from AI, and only 4% were creating substantial value (BCG, Oct 2024). McKinsey tested 25 attributes and found workflow redesign had the biggest effect on EBIT impact — and only 21% of companies had done it (McKinsey, Mar 2025). Gartner predicts 30%+ of generative-AI projects abandoned after proof of concept and 40%+ of agentic projects cancelled by 2027 (Gartner 2024, Gartner 2025). Predictions, not outcomes. Gartner’s track record on these is mixed. Still, nobody is predicting the opposite.

It is not only an engineering story. Typeface surveyed 200+ marketing VPs in 2026: 93% said AI had raised the pressure to move faster, while the share needing one to two months to launch a campaign rose from 5% to 34% in a year (Typeface Signal Report). Typeface sells the tools that supposedly fix this, so read the numbers with that in mind. But the shape is familiar. Creation got fast; approvals, hand-offs and governance did not.

BCG’s rule of thumb is the one we keep coming back to: the companies that get value from AI put roughly 10% of their effort into algorithms, 20% into data and technology, and 70% into people, process and change (BCG, Jan 2025). It is a heuristic, not a measurement. It is also the exact inverse of how most AI-built MVPs are put together — all algorithm and interface, no operating model. Ours included, at the start.

Where do AI-built products stall between demo and production?

The demo works. What is missing is everything a demo never has to survive.

The five stall patterns: invisible security debt, no one on call, integration drift, pilot purgatory, governance at the end. The first three are technical readiness; the last two are organizational.

Invisible security debt is the first, and it is the one AI is worst at hiding. Keys in code or in git history. Authorization checked in the browser and nowhere else. Database rules left permissive because the demo needed them that way. Dependencies nobody looked at. The cost is a leak or a breach that ends the product before it starts, and the awkward truth is that the model that wrote the code will cheerfully tell you it is secure.

No one is on call. No error tracking, no alerts, backups that have never been restored, no rollback, no runbook. The first outage lasts days because nobody knows it is happening. The first data loss is permanent because the backup was a checkbox.

Integration drift is the quiet one. Sandbox keys still in production. API versions unpinned. Webhooks accepted without verification. No plan for the day a vendor has an outage or changes its terms. It shows up as silent failures and surprise bills from systems you do not own.

Those three are technical readiness, and any decent checklist covers them. The next two are where we see products actually die.

Pilot purgatory. No agreed success metric, no “scale or kill” criteria, the new tool launched while the old workflow stays in place. Months of “promising” with no decision, and users drifting back to the way they worked before.

Governance at the end. Legal, security and brand review as one final gate, with everything queued behind it. The fastest build meets the slowest sign-off, and the speed you gained in week one disappears in month four.

Most checklists cover the first three. Most failures come from the last two.

What does a production readiness framework for AI products look like?

Readiness is a profile, not a checklist. That distinction took us a while to arrive at. A checklist treats a weekend prototype and a regulated B2B product the same way, so it is either too heavy for one or too light for the other, and either way people stop using it.

The Viaknox readiness framework: 13 core checks (10 technical, 3 organizational), 10 modules switched on by profile, 3 tiers that set the bar.

Our version has three parts. Thirteen core domains apply to every product: ten technical (secrets and config, identity and access, edge and app security, data and backups, reliability, observability, CI/CD and release, testing, performance and cost, operations and support) and three organisational (legal and trust, product readiness, adoption and operating model). Ten modules switch on depending on what you are building — integrations consumed, public API, mobile and app stores, web and email, AI and LLM features, payments, regulated data, B2B and enterprise, marketplaces, dev tools and open source. A mobile AI app that takes payments turns on three of them. A developer plugin turns on two different ones.

Then the tier sets how high the bar is, and this is what keeps the framework honest in both directions. Enterprise controls on a fifty-user beta waste months. Beta controls on a paid product lose customers.

  • T1, public beta. Real users, low scale, some downtime tolerated. No leaked secrets, solid auth, backups that restore, errors tracked, legal pages, a rollback that works.
  • T2, general availability. Paying users with SLA expectations. Everything in T1 plus service targets, alerts that reach a person, load testing, runbooks, a support path, fallbacks for every integration.
  • T3, enterprise or regulated. B2B contracts, health, payment or children’s data. Everything in T2 plus SSO, audit logs, data agreements, a pen test, SOC 2 or HIPAA controls, disaster-recovery drills.

Three rules make it work, and the first one is the one we had to learn the hard way.

No evidence, no pass. Every item is Pass, Gap, N/A or Unknown, and a Pass needs proof — a file, a config, a test run, a dashboard. A README that says “backups are enabled” is not proof. Neither is the coding agent saying so. We had agents mark items complete that a five-minute check showed were not, which is how this rule got its name.

Verify the rules live. App store, OAuth, payment and email-sender requirements change every year. Check the vendor’s current documentation, not the model’s memory of it.

Governance runs inside the workflow. The checks live in CI, in pull requests and in the product itself, not in a meeting the week before launch.

Readiness Check · Step 1 of 3

How ready is your AI-built product? Ten questions, no evidence required — yet.

Answer honestly. A "Yes" you can't prove is a "No" in production.

  1. 01 SecretsHave you scanned your full git history for leaked keys and rotated anything found?
  2. 02 IdentityIs authorization enforced on the server for every data write, not just in the UI?
  3. 04 DataHave you restored a backup end-to-end at least once, and timed it?
  4. 05/06 Reliability & observabilityIf your product failed at 3 a.m., would a human be alerted, and could they roll back?
  5. 07 CI/CDDoes main require lint, tests, a secret scan and a dependency audit to pass?
  6. 09 CostIs there a hard spend cap on every paid API you call (LLM, SMS, email, cloud)?
  7. 11 LegalDo your Terms and Privacy Policy describe the data you actually collect and the AI output you show?
  8. 13 AdoptionIs there one written success metric with a target date — and an agreed "scale or kill" rule?
  9. 13 AdoptionIs it written down who approves a low-risk change vs a high-risk one, and how fast?
  10. 13 AdoptionIs there a named owner for this product after launch?

Why does a product that passes every technical check still fail?

Because nobody changed how the work gets done around it. Shipped is not adopted.

The fix is six decisions made before launch rather than after, and none of them are technical. What single number makes this launch a win, and by when — without it, every pilot looks promising forever. What result means scale it, and what result means stop. Who approves what, and how fast — low-risk changes get one reviewer, high-risk ones get more, and the turnaround is published. Whether this touches personal data, AI output or regulated users, decided on day one rather than discovered in month three. Which early users will try it first and teach the others, because peers pull adoption faster than mandates. And when usage gets reviewed — 30, 60, 90 days, then quarterly — so the product is treated as something alive rather than something released.

The approval decision is the one most teams underestimate. In Typeface’s survey, 88% of marketing leaders said content creation was fast and sign-off was not. Compliance was the top barrier to scaling AI, cited by 66% (Typeface). Same vendor caveat as before, but we would not argue with either number.

Two design choices help adoption stick. Put the capability inside the tools people already use — their editor, CMS, inbox or IDE — instead of a new app they have to remember to visit. And show the limits: what the AI cannot do, where the answer came from, which parts are drafts. Trust that is earned holds; trust that is assumed does not.

This is where BCG’s 70% goes, and McKinsey’s data says the same thing from the other side: redesigning the workflow is the single biggest driver of bottom-line impact from generative AI (McKinsey).

How do you get an AI-built product production-ready in 30 days?

Same framework, three entry points, depending on who you are.

Founders and product leaders

The 30-day path to Tier 1: week 1 profile and decide, week 2 close the blockers, week 3 make it operable, week 4 launch to a small group.

Week one is profile and decide: tier, success metric, exit criteria, every integration listed, legal and privacy flagged. Done when there is one page everyone has signed and the readiness scan has run and logged its gaps. Week two is closing the blockers — secrets out of code and out of history, authorization on the server, a backup actually restored once, error tracking live — done when there are zero open blockers for T1 and each fix has evidence attached. Week three makes it operable: rollback proven, alerts that reach a human, a runbook for the top five failures, spend caps on every paid API. Done when you have staged a failure and watched it get detected, alerted and rolled back. Week four is launch to a small group, champions onboarded, a feedback channel open, the first review date on the calendar.

The founder’s job is not to do each check. It is to refuse to launch on “we think it’s fine.”

Enterprise leaders

Put governance inside the workflow, once. Security, brand, privacy and approval rules become checks in CI, pull-request templates and product guardrails rather than a review meeting. Tier the approvals by risk and publish the turnaround for each tier. Standardise before you scale — only 20% of marketing organisations have documented workflows ready for AI (Typeface), and in our experience software teams are not much better, so one shared readiness profile beats a different checklist per team. And plan the exit: own your data, prompts and configuration, avoid multi-year lock-in without break clauses, and review models and vendors twice a year, because the one you chose this spring will not be the best one next spring.

Engineering leads

Make every check produce evidence. Required checks on main: lint, types, tests on the critical paths, a secret scan over the full history, a dependency audit. Proof over promises — a restore log, a rollback run, an alert that fired, a load-test report, each one attached to the readiness item it closes. Test the critical paths before chasing a coverage number: auth, payments, permissions and data writes first, because a percentage is easy to game and a broken permission check is not. And verify every template an agent hands you. AI-suggested CI files routinely reference actions and packages that do not exist. Check that each one resolves before you trust it.

We’re running it on our own product first

It would be easy to publish a framework and never use it. So the first product through it is ours, and we are going to publish the results whether they flatter us or not.

KnoKeep by Viaknox is portable, secret-safe memory that lets an AI coding agent resume exactly where the last session stopped, even in a different tool. It ships as a Claude Code plugin, a Cursor rule, a Codex instruction file and a shell command. It is open source, AGPL-3.0, at v0.1.0, with a commercial licence planned for teams that embed it in their own products.

On paper it is in good shape: 683 tests passing, CI on Linux, macOS and Windows, gitleaks on every commit, a threat model, branch protection, and a review step where several different models argue with each other about every change before a human sees it. A technical-only checklist would wave it through. That is exactly why it is a useful test.

Its profile is T1 — a public open-source beta, with T2 in view for the commercial licence — with the Integrations, Marketplaces and Dev-tools modules switched on. The questions we expect it to struggle with are not the technical ones. What is the upgrade path when the stored memory format changes? What does AGPL actually mean for a company that wants to use it internally? Does a clean-machine install work on all four tools, every time, or only on the machines we built it on? And what is the success metric for the beta — because right now, if we are honest, we have a feeling rather than a number.

The full scorecard, the blockers we find, what we fix and how long it takes will be the next post. No evidence, no pass applies to us too.

The free Production Readiness Kit

The framework in this article is free, and it works with the AI coding tool you already use. It is built on the open SKILL.md format that Claude, Codex, Cursor, GitHub Copilot in VS Code and Gemini CLI can all load (VS Code docs).

What’s in the kitWho uses itHow
Readiness skill (SKILL.md)Engineers, via their AI coding agentAsk the agent to “run production readiness”; it profiles the repo, gathers evidence and returns a scorecard and a gap backlog
Tool adaptersTeams on Cursor, Codex or plain shellsA Cursor rule and an AGENTS.md that point each tool at the same skill
Printable scorecardFounders, product and enterprise leadersThe 13 domains and 10 modules as a one-page self-check, no code needed
Gap ticket templateAnyone running the backlogEach gap as a ticket: why, done-when, evidence to close, tier, severity, size

License: MIT for the code, CC BY 4.0 for the documents, so you can adapt it for your own team.

→ Get the kit: github.com/viaknox/production-readiness

Not ready to install anything? The ten-question Readiness Check above gives you a score, your weakest domains and the install command for your tool in under two minutes.

Readiness Check · Step 1 of 3

Score yourself in ten questions.

Show the ten questions (about two minutes)
  1. 01 SecretsHave you scanned your full git history for leaked keys and rotated anything found?
  2. 02 IdentityIs authorization enforced on the server for every data write, not just in the UI?
  3. 04 DataHave you restored a backup end-to-end at least once, and timed it?
  4. 05/06 Reliability & observabilityIf your product failed at 3 a.m., would a human be alerted, and could they roll back?
  5. 07 CI/CDDoes main require lint, tests, a secret scan and a dependency audit to pass?
  6. 09 CostIs there a hard spend cap on every paid API you call (LLM, SMS, email, cloud)?
  7. 11 LegalDo your Terms and Privacy Policy describe the data you actually collect and the AI output you show?
  8. 13 AdoptionIs there one written success metric with a target date — and an agreed "scale or kill" rule?
  9. 13 AdoptionIs it written down who approves a low-risk change vs a high-risk one, and how fast?
  10. 13 AdoptionIs there a named owner for this product after launch?

Sources and notes on the evidence

SourceUsed forCaveat
DORA, Accelerate State of DevOps 2024Stability −7.2%, throughput −1.5% per 25% more AI adoptionModelled system-level estimate, not a measured drop per team
DORA, State of AI-assisted Software Development 2025“AI as amplifier”—
METR, Jul 202519% slower; believed 20% fasterRandomized study of 16 experienced open-source developers on their own repositories
Veracode, 2025 GenAI Code Security ReportSecurity flaws in 45% of testsVendor test suite across 100+ models, not production code
Stack Overflow Developer Survey 202546% distrust vs 33% trust—
BCG, Where’s the Value in AI?, Oct 202474% no tangible value; 4% creating substantial valueThe 4% and the 22% who show some value are different groups; don’t merge them
BCG, From Potential to Profit: Closing the AI Impact Gap, Jan 2025The 10-20-70 ruleDescribes what AI leaders do; a heuristic, not a measured allocation
McKinsey, The State of AI: How organizations are rewiring to capture value, Mar 2025Workflow redesign is the top driver of EBIT impact; 21% have done itFigures from the March 2025 edition; later editions differ
Gartner, Jul 2024 and Jun 202530%+ abandoned after proof of concept; 40%+ agentic canceled by 2027Predictions, not measured outcomes
Typeface, Signal Report 04: The AI Speed Paradox, 2026Campaign timelines, sign-off bottleneck, 20% documented workflows, 66% compliance barrierVendor survey of 200+ marketing VPs; the conclusions align with what the vendor sells

Questions about this article

Why do AI-built products take so long to ship if they are so fast to build?

Because the demo is the easy part. A demo never has to survive a leaked key, a 3 a.m. outage, a vendor API change, an unagreed success metric or a legal review. Those five things — not the model — are where AI-built products stall, and none of them get faster because the code was written by an AI.

What is a production readiness framework for AI products?

A fixed set of checks every product must pass before real users depend on it, scored with evidence. Viaknox's version has 13 core domains (10 technical, 3 organizational), 10 situational modules switched on by what you are building, and three tiers that set how high the bar is. The rule is simple: no evidence, no pass.

How long does it take to get an AI-built MVP production-ready?

For a public beta (Tier 1) with a small team, about 30 days if you work the checks in order: profile and decide, close the blockers, make it operable, launch to a small group. Enterprise or regulated tiers take longer because they add SSO, audit logs, pen tests and data agreements.

Is the Viaknox Production Readiness Kit free?

Yes. The skill, tool adapters, printable scorecard and gap-ticket template are MIT (code) and CC BY 4.0 (documents). It runs inside Claude, Cursor, Codex, GitHub Copilot or Gemini CLI, or as a one-page self-check with no code at all.

Sources

  1. DORA — Accelerate State of DevOps Report 2024
  2. DORA — State of AI-assisted Software Development 2025
  3. METR — Measuring the impact of early-2025 AI on experienced open-source developer productivity
  4. Veracode — 2025 GenAI Code Security Report
  5. Stack Overflow — 2025 Developer Survey: AI
  6. BCG — Where's the Value in AI? (Oct 2024)
  7. BCG — From Potential to Profit: Closing the AI Impact Gap (Jan 2025)
  8. McKinsey — The State of AI (Mar 2025)
  9. Gartner — 30% of GenAI projects abandoned after PoC (Jul 2024)
  10. Gartner — 40% of agentic AI projects canceled by 2027 (Jun 2025)
  11. Typeface — Signal Report 04: The AI Speed Paradox (2026)
  12. VS Code docs — Agent Skills

Have a decision or a system to get right?

Consulting: Assess, Architect, Deliver — with your team, not around it.

How an engagement works →

Tags: production-readinessai-productsshippinggovernancefounders