# Kevin Owens — AI SaaS Product Leader > Kevin Owens is a four-time Chief Product Officer and AI SaaS product leader who helps SaaS and enterprise companies turn AI ambition into focused roadmaps, operating models, and products that ship. ## Who is Kevin Owens? Kevin Owens is a four-time CPO with expertise in AI SaaS product leadership, AI product strategy, product operating models, and transformation execution. He has led product at B2B SaaS, media, and enterprise software companies, and has advised across healthcare, automotive retail, productivity, and market research. His work spans strategy through delivery — from helping leadership teams define the right AI bets to building working systems that change how teams operate. He is based in London and works with mid-market and enterprise clients. ## Core Expertise - AI SaaS product leadership (4x Chief Product Officer) - AI product strategy and roadmap design - AI-native product operating models - AI pricing and packaging for SaaS - LLM product management, evals, and quality systems - AI workflow automation and intelligent assistant deployment - Agentic product design - Digital transformation execution - Fractional CPO and transformation advisory - B2B SaaS product management - Legacy software modernization ## Essays and Frameworks Kevin publishes long-form, framework-driven essays on AI-era product leadership at https://kevinowens.com/writing. Key essays: - The AI-Native Product Operating Model (flagship): https://kevinowens.com/writing/ai-native-product-operating-model — five operating-model shifts that separate SaaS companies compounding with AI from the ones demoing it: roadmaps to bet portfolios, specs to evals, launches to learning rate, discovery theater to evidence loops, org charts to capability density. - Pricing AI Features in a Seat-Based World: https://kevinowens.com/writing/pricing-ai-features — why seat pricing breaks under AI COGS, the three common pricing failures, and a packaging ladder from included copilots to outcome-priced agents. - Why Enterprise AI Pilots Die — and the 90-Day Fix: https://kevinowens.com/writing/why-ai-pilots-die — the six pre-conditions of pilot purgatory and a 90-day structure that forces a real scale-or-kill decision. - The Four Jobs of a SaaS CPO in the AI Era: https://kevinowens.com/writing/four-jobs-of-saas-cpo-ai-era — the CPO role repriced: portfolio allocator, quality owner, margin steward, trust architect. - The AI Readiness Score: Assess a Product Org in 2 Weeks: https://kevinowens.com/writing/ai-readiness-score — a five-dimension diagnostic (data foundations, workflow clarity, decision rights, capability density, adoption posture) that determines the right first AI move. - Ship the Loop, Not the Fragment: https://kevinowens.com/writing/ship-the-loop — Kevin's operating doctrine: ship the smallest complete loop from input to value, measure what it changes, keep what compounds, kill what doesn't. ## Services Kevin's work is organized across three practices: ### ProductExec (https://productexec.co) Executive advisory for AI product strategy. Ideal for companies that need to turn AI pressure into product strategy, operating model decisions, and an executable roadmap before teams commit delivery capacity. Key services: - AI Product Strategy Diagnostic - Product Operating Model Reset - AI Product Strategy Sprint - Fractional CPO / Transformation Advisor Contact: https://productexec.co/contact ### Enterprise AI Studio (https://enterpriseaistudio.com) AI implementation services for mid-market and enterprise teams. Builds working AI systems — from pilots to production. Key services: - AI Transformation Readiness Audit - AI Implementation Pilot (automation, assistants, document intelligence, workflow agents) - Legacy-to-AI Modernization Plan - Fractional AI Engineering Contact: https://enterpriseaistudio.com ### Smart Biz AI Hub (https://smartbizaihub.com) Practical AI pilots and guides for smaller teams — working pilots measured at 30, 60, and 90 days. ## Professional Links - Website: https://kevinowens.com - Essays and frameworks: https://kevinowens.com/writing - Speaking: https://kevinowens.com/speaking - Representative work: https://kevinowens.com/work - Advisory practice: https://productexec.co - Implementation practice: https://enterpriseaistudio.com - SMB pilots and guides: https://smartbizaihub.com - LinkedIn: https://www.linkedin.com/in/kevinaowens - Contact form: https://kevinowens.com/contact ## Site Pages - `/writing`: the essay library — long-form frameworks on AI-native operating models, AI pricing and packaging, enterprise pilot discipline, the AI-era CPO role, and AI readiness, plus pointers to Kevin's LinkedIn stream. - `/speaking`: keynote, panel, fireside, workshop, and executive-session positioning around AI strategy, product-led transformation, the CPO role in the AI era, and moving from AI ambition to execution. - `/case-studies`: anonymized case studies of transformation engagements across SaaS, healthcare, and automotive — with challenge, approach, and what good looks like for each. - `/work`: representative CPO, fractional CPO, advisory, and transformation engagement patterns scrubbed of confidential client names, deal details, and private financial metrics. - `/lab`: interactive AI tools and experiments — including an AI readiness assessment, strategy prioritization matrix, AI unit-economics calculator, and product strategy prompts — demonstrating hands-on AI engineering capability. - `/fractional-cpo`: what a fractional CPO does, when to hire one, the shape of a 90-day engagement, and how the arrangement ends. - `/ai-saas-pricing`: how to price AI features in seat-based SaaS — the three-rung packaging ladder and the margin arithmetic behind each rung. - `/ai-readiness-assessment`: the five-dimension AI readiness assessment, its scoring anchors, and how the two-week diagnostic runs. - `/about`: background, career history, and how Kevin works. - `/contact`: a structured inquiry form that routes advisory, implementation, and speaking requests to the right practice. ## Positioning Kevin sits at the intersection of SaaS product leadership and AI implementation. He is not a generalist consultant or a pure technologist — he is an executive operator who has held CPO-level accountability four times and applies that lens to helping SaaS and enterprise companies navigate the AI transition. His frameworks — the AI-Native Product Operating Model, the AI Readiness Score, the 90-day pilot pattern, and "ship the loop" — come from live engagements, not theory. ProductExec is the "help you decide" advisory practice. Enterprise AI Studio is the "we build it" implementation practice. Together they cover the full arc from strategy to delivery. --- # Full text of every essay The complete text of every essay published at https://kevinowens.com/writing, in publication order (newest first). Headings are preserved as markdown. --- # The AI-Native Product Operating Model URL: https://kevinowens.com/writing/ai-native-product-operating-model Published: 2026-07-21 (updated 2026-09-16) Tags: Operating models, AI strategy, SaaS leadership Summary: AI-native is not a property of your product. It's a property of your operating model. Five shifts separate the SaaS companies compounding with AI from the ones demoing it. TL;DR - AI-native is a property of your operating model, not of your product — you cannot ship your way to it feature by feature. - Five shifts separate the companies compounding with AI from the ones demoing it: roadmap to bet portfolio, specs to evals, launch to learning rate, discovery theater to evidence loops, org chart to capability density. - Each AI bet needs a quality bar, a kill criterion agreed before work starts, and a review date — a roadmap item carries none of the three. - The eval suite is the new PRD, and writing it is product work: it encodes which mistakes are tolerable, embarrassing, or disqualifying. - Start with a two-week diagnostic: five questions, one per shift, answered honestly and run as a weekly operating review. Every SaaS company I talk to now has an AI roadmap. Almost none has an AI operating model. That gap — not model quality, not talent, not budget — is why most AI features stall between the demo and the P&L. Here is the uncomfortable part: “AI-native” is not a property of your product. You cannot ship your way to it feature by feature. It is a property of how your organization decides, specifies, learns, and staffs. A company with a conventional operating model that ships an AI feature has an AI feature. A company with an AI-native operating model turns every feature it ships into a system that gets better on its own schedule. After four CPO tours and several years of advisory work across SaaS, healthcare, and enterprise software, I see the same five shifts in every organization that makes the transition — and the same absences in every organization that doesn't. ## Shift 1: From roadmap to bet portfolio A classic SaaS roadmap is a promise ledger: features, dates, owners. It assumes the work is deterministic — that effort in produces feature out. AI work is probabilistic. Some bets will not clear the quality bar no matter how much effort you spend, and you often cannot know which until you are three weeks in. Running probabilistic work through a deterministic roadmap produces two failure modes: - Sandbagging. Every AI item becomes “exploration,” and nothing ships. - False precision. Leadership treats model behavior as an engineering estimate problem and burns trust when the date slips for the third time. The fix is to run AI investment as a portfolio with explicit horizons: - Core bets. They harden workflows you already win with. - Adjacent bets. They automate a step your customers do around your product today. - Frontier bets. They would change your category if the capability curve cooperates. Each bet carries three things a roadmap item never does: - A quality bar it must clear. - A kill criterion agreed before work starts — the same discipline that separates pilots that graduate from pilots that never die. - A review date when the bet is re-priced or retired. A roadmap asks “when will it ship?” A portfolio asks “what would make us stop?” The second question is the one that protects your margin and your credibility. ## Shift 2: From specs to evals In deterministic software, the spec is the contract: acceptance criteria, edge cases, done means done. For AI behavior, prose specs are unfalsifiable. “The assistant should answer billing questions accurately and escalate when unsure” sounds like a requirement. It is actually a wish. The AI-native replacement is the eval: a graded set of real examples that defines what good looks like, executable on every model change, prompt change, and retrieval change. The eval suite is the new PRD. Writing it is product work, not engineering work — it encodes judgment about which mistakes are tolerable, which are embarrassing, and which are disqualifying. That judgment is precisely what a product organization exists to hold. A practical bar I give teams: no AI behavior ships without an eval that a new PM could run in an afternoon and interpret without asking anyone. If your team cannot say what score the feature gets today and what score it needs, you do not have a quality bar. You have vibes. ## Shift 3: From launch to learning rate SaaS product culture is launch culture: gate, announce, move on. AI product value is delivered on a different clock. The feature at launch is the worst version customers will ever use — if, and only if, you built the loop that improves it: instrumentation on real usage, feedback capture in the workflow, an eval that turns feedback into the next quality target, and a cadence that ships the improvement. So the metric that matters is not features per quarter. It is learning rate: how much does the system improve per week of real usage? Two teams can ship the same feature on the same day; six months later one has a compounding asset and the other has a stale demo. The difference was never the model. It was whether anyone owned the loop after launch week. This is the discipline I compress into three sentences: ship the loop, not the fragment. Measure what it changes, not what it demos. Keep what compounds — and kill what doesn't. ## Shift 4: From discovery theater to evidence loops Conventional discovery answers “do customers want this?” AI discovery has to answer a harder compound question: “does the capability clear the quality bar, in this workflow, at a cost that survives our pricing?” A promising interview plus an impressive demo answers neither. Evidence loops replace discovery phases. Take the narrowest real slice of the workflow, run it with production-shaped data behind a human checkpoint, and measure three numbers from week one: - Quality against the eval. The same suite you would ship against, run on real slices rather than the demo path. - Cost per completed task. The number that decides whether the feature survives your pricing. - Human acceptance. Whether the people in the loop actually kept the output. Those three numbers — not stakeholder enthusiasm — decide whether the bet graduates. The teams that struggle most here are the ones proudest of their discovery practice. Their rituals generate conviction. AI work punishes conviction that arrives ahead of evidence. ## Shift 5: From org chart to capability density The instinctive organizational response to AI is additive: an AI team, an AI PM title, a center of excellence. This quarantines the capability exactly where it cannot compound. The AI team becomes a service desk with a backlog; every product team's AI idea enters a queue owned by someone else; nothing changes about how the core organization works. The AI-native pattern is density, not specialization: every product trio gains working fluency — PMs who can read an eval and reason about failure modes, designers who treat model uncertainty as a design material, engineers who can wire a retrieval pipeline without a platform ticket. A small enablement group can seed this. It cannot substitute for it. The test I use: pick a random product team and ask them to explain their feature's worst failure mode and what it costs per thousand requests. If the answer is “we'd have to ask the AI team,” you have an org-chart problem wearing a technology costume. The five shifts of the AI-native product operating model ShiftFromToThe artifact that proves it 1RoadmapBet portfolioKill criterion written before work starts 2Prose specsEvalsAn eval a new PM can run in an afternoon 3LaunchLearning rateMeasured improvement over the last 30 days 4Discovery theaterEvidence loopsCost per completed task on real usage 5Org chartCapability densityA random team explaining its failure modes unaided ## What does not change Every operating-model conversation eventually produces someone arguing that AI changes everything. It does not. Strategy is still choice under constraint. Distribution still beats features. Customer trust is still earned in years and lost in incidents. Unit economics still decide who survives a downturn. The five shifts above change how you execute; they do not repeal why you exist. If anything, AI raises the price of strategic sloppiness, because it hands every competitor the same capability curve and rewards the organization that metabolizes it fastest. ## Where to start: the two-week diagnostic You do not need a transformation program to begin. You need honest answers to five questions, one per shift. (The fuller organizational version of this diagnostic is the AI Readiness Score.) - Portfolio: Can you list your AI bets with the kill criterion attached to each? Written down, before the work started? - Evals: For your most visible AI feature, what is its eval score today — and what score does it need? - Learning rate: Is that feature measurably better than it was 30 days ago? Who owns making that true? - Evidence: What is your cost per completed task on real usage — not the demo path? - Density: Can a randomly chosen product team explain their feature's failure modes without escalating? Most leadership teams can answer one of the five. That is not a criticism — it is a baseline. The organizations that will own their categories in three years are the ones treating these five answers as the operating review, run weekly, starting now. ### FAQ **What is an AI-native product operating model?** It is the way an organization decides, specifies, learns, and staffs around probabilistic software — not a property of the product itself. In practice it means running AI work as a bet portfolio, defining quality with evals instead of prose specs, managing learning rate after launch, and spreading fluency rather than quarantining it. **Why do AI roadmaps stall between demo and P&L?** Because probabilistic work is being run through a deterministic roadmap. Teams either sandbag every AI item into exploration, or leadership treats model behavior as an estimate problem and burns trust when dates slip. The fix is the portfolio: every bet carries a quality bar, a kill criterion set before work starts, and a review date. **How do I know if my company is AI-native yet?** Run the two-week diagnostic: five questions, one per shift. Can you list your AI bets with kill criteria? What is your flagship feature's eval score, and what does it need? Is it better than 30 days ago? What is your cost per completed task on real usage? Can a random team explain its failure modes unaided? Most leadership teams can answer one of the five. --- # Pricing AI Features in a Seat-Based World URL: https://kevinowens.com/writing/pricing-ai-features Published: 2026-06-30 (updated 2026-09-16) Tags: Pricing & packaging, SaaS economics, AI strategy Summary: Seat pricing was built on an assumption AI just broke: that serving one more user costs you nothing. How SaaS leaders should price intelligence without eroding margin or suppressing adoption. TL;DR - Seat pricing assumes a near-zero marginal cost per user; AI features carry a real unit cost that scales with usage, so your most engaged accounts become your most expensive. - The three common failures are bundling AI into existing seats, metering everything, and packaging it all into one detachable AI SKU. - Price on a three-rung ladder that follows how much of the job the AI completes: assistive included in seats, workflow as platform fee plus metered bands, outcome priced as digital labor. - Product — not finance — owns three numbers per feature monthly: cost per completed task, gross margin by rung and segment, and cost trajectory. - Credits are a transitional currency for learning the real unit of work, not a pricing model to keep past year two. Seat-based SaaS pricing rests on one quiet assumption: marginal cost is approximately zero. One more user, one more login, no meaningful change to your cost of goods. Twenty years of SaaS economics — the 80% gross margins, the land-and-expand playbook, the per-seat upsell — sit on top of that assumption. AI features break it. Every generation, every retrieval, every agent step has a real unit cost, and that cost scales with usage, not with seats. Your most engaged customers — the ones every SaaS playbook tells you to celebrate — become your most expensive to serve. Pricing is no longer a go-to-market decision that product hears about later. It is a product architecture decision, and product leaders who treat it as someone else's spreadsheet will ship features that lose money at their moment of greatest success. ## The three ways SaaS teams get it wrong - Bundling AI into existing seats. The path of least resistance: fold the AI features into current plans, call it product velocity, hope the retention lift covers the cost. Sometimes it does — briefly. Then a segment of power users discovers the feature, usage follows a power law (it always follows a power law), and your gross margin quietly bleeds out through the top decile of accounts. Worse, you have now taught the market that intelligence is free, and the first repricing conversation arrives after the anchor has set. - Pure usage metering. The opposite reflex: meter everything, pass the cost through, stay safe. Economically clean, behaviorally poisonous. A visible meter makes every use of the feature a small purchasing decision, and users respond the way they respond to all metered utilities — they conserve. Adoption stalls precisely where you need habit formation. You protected your margin on a feature nobody learned to rely on. - The bolt-on AI SKU. Packaging every AI capability into a single “AI add-on” tier feels tidy and demos well in a pricing deck. But it fragments value: workflow features that belong in the core product get held hostage to an upsell, sales teams discount the SKU to close, and the capabilities inside it never integrate deeply because they must remain detachable. An add-on is a packaging decision masquerading as a strategy. Price the outcome where you can, the workflow where you must, and the token never. Customers don't buy inference. They buy finished work. ## A packaging ladder that follows value The pattern I recommend to SaaS leadership teams is a ladder with three rungs, each matching how much of the job the AI actually completes: - Assistive (included in seats). Copilot-style features that make a human faster inside your existing workflow — drafting, summarizing, suggesting. Include them. Their job is retention and differentiation, their cost per seat is bounded because a human is the throttle, and withholding them just invites a competitor to make them table stakes. Watch the margin, set generous-but-real fair-use ceilings, and move on. - Workflow (hybrid: platform fee + metered band). Features that complete a defined unit of work with light supervision — processing a document, triaging a ticket, generating a report. Price them per unit of work in bands: a platform fee that covers a healthy allotment, then predictable tiers. The unit must be one the customer already counts — documents, tickets, campaigns — never tokens or credits they have to learn. - Outcome (agent pricing). Agents that own a job end to end are not features; they are digital labor, and the comparison price in the buyer's head is a salary or a BPO contract, not a software line item. Price against the outcome — per resolved case, per qualified lead, per closed book month — with quality-linked terms. This is the rung where the margin structure of your company actually changes, and it deserves CFO-level partnership, not a pricing-page tweak. The packaging ladder: how much of the job the AI completes, and how to price it RungWhat the AI doesPricing modelUnitMargin risk AssistiveMakes a human faster inside the existing workflowIncluded in the seat, with fair-use ceilingsThe seatBounded — the human is the throttle WorkflowCompletes a defined unit of work with light supervisionPlatform fee plus metered bandsA unit the customer already countsReal — concentrated in the top decile of accounts OutcomeOwns a job end to end as digital laborOutcome pricing with quality-linked termsThe resolved case, lead or closed monthStructural — changes the company's margin profile The rung a feature belongs on is a question about capability, not ambition, and the honest answer usually comes from an eval rather than a demo — the same discipline I argue for in shifting from specs to evals. A feature you cannot measure completing the job is not an outcome-priced feature yet. ## The margin discipline behind the ladder None of the ladder works unless someone owns the unit economics at the workflow level, and in an AI-era SaaS company that someone is product — the margin steward job. Three numbers per AI feature, reviewed monthly: - Cost per completed task — fully loaded: inference, retrieval, retries, the failure paths. Demo-path cost is fiction; power-law users and retry storms are where the money goes. It is the same number that decides whether a loop is worth keeping in keep what compounds — and kill what doesn't. - Gross margin at the current price point — per rung, per segment. “Blended margin looks fine” is how the top decile eats you. - Cost trajectory — model prices fall, but your usage mix shifts toward heavier tasks as trust grows. Model both curves; the second one is usually steeper. A note on credits, because every pricing meeting eventually proposes them: credits are a transitional currency, useful when your unit of work is still unstable. They buy flexibility at the cost of comprehension — every credit balance is a small anxiety generator, and customers who cannot predict their bill discount the product to compensate. Use credits to learn the real unit of work, then retire them into the ladder. A credit system that persists past year two is a pricing decision you are refusing to make — and it deserves the same kind of stated kill criterion that keeps a pilot from becoming permanent in killing pilots is a capability. ## Five questions before your next AI pricing decision - Unit the customer counts. What unit does the customer already count that this feature maps to? - Cost at the tail, not the median. What is the fully loaded cost per completed task at the 90th-percentile account? - Honest rung. Which rung of the ladder is this — and are we pricing it on that rung, or on the rung that is easier to sell? - Margin under 10x usage. If usage grows 10x in our best accounts, does gross margin hold above the line the CFO signed up for? - The anchor you are setting. What price anchor are we setting for the agentic version of this feature two years out? Seat pricing is not dead — assistive AI arguably strengthens it. But the era when product could ship value and let finance find the price is over. In AI-era SaaS, the pricing model is part of the product architecture. Leaders who internalize that will fund their own compounding. Leaders who don't will discover their most successful feature is also their most expensive mistake. ### FAQ **Should AI features be included in the seat price or charged separately?** Both, depending on the rung. Assistive copilots that only make a human faster belong in the seat price, because a human throttles the cost. Work the AI completes on its own should carry its own price: a platform fee plus metered bands for workflow features, and outcome pricing for agents. **How do you price AI features without killing adoption?** Avoid a visible meter on habit-forming features. Users conserve whatever they can see ticking, so adoption stalls exactly where you need repetition. Include assistive features, and price workflow features in a unit the customer already counts (documents, tickets, cases) rather than tokens or credits. **What metrics should product leaders track for AI feature margin?** Three, reviewed monthly per feature: fully loaded cost per completed task including retries and failure paths, gross margin per rung and per segment rather than blended, and cost trajectory as the usage mix shifts toward heavier tasks. --- # Why Enterprise AI Pilots Die — and the 90-Day Fix URL: https://kevinowens.com/writing/why-ai-pilots-die Published: 2026-06-09 (updated 2026-09-16) Tags: AI transformation, Execution, Enterprise AI Summary: Most enterprise AI pilots don't fail. They just never end. The six reasons pilots stall in purgatory, and the 90-day structure that forces a real scale-or-kill decision. TL;DR - Enterprise AI pilots rarely fail outright; they stall in purgatory, consuming budget and credibility while proving nothing. - Six pre-conditions cause the stall: no quality bar, no baseline, a pilot beside the workflow, no operator owner, undesigned adoption, and success defined as sentiment. - The 90-day pattern answers all six with four gated phases: baseline and definition of done, a narrow wedge in production, instrumentation of quality, cost and trust, then a forced decision. - Human acceptance rate — how often the checkpoint takes the output unchanged — is the curve teams skip and the one that predicts whether a pilot scales. - A clean kill on the pre-committed date is a capability, not a failure: it prices the capability curve and says what must become true before the next attempt. The most common state of an enterprise AI initiative is not success or failure. It is purgatory: a pilot that demos well, offends no one, proves nothing, and never ends. Six months in, nobody can say whether it worked, because nobody defined what working meant. The pilot is not dying. It is worse — it is undead, consuming budget and credibility while blocking the honest conversation a failed experiment would have forced. I have been brought in behind enough of these to see the pattern clearly. Pilots do not stall for technical reasons. They stall for product-management reasons that were locked in before the first line of code. ## The six pre-conditions of pilot purgatory - Nobody answered “how good is good enough.” Every stalled pilot shares this. The team can tell you the model, the vendor, the architecture — but not the accuracy threshold at which the business would act. Without a pre-committed quality bar, every result is interpretable, and interpretable results extend pilots forever. The bar is a business decision, not a data-science one: what error rate would you accept from a new hire doing this task? It is the same bar an eval exists to measure. - There is no baseline. “The AI resolves tickets in four minutes” means nothing if nobody measured the human process first. Teams skip baselining because it is boring and slows the kickoff. Then they spend month five arguing about whether four minutes is good — an argument that costs more than the two weeks of measurement would have. - The pilot runs beside the workflow, not inside it. The sandbox pilot — a parallel tool, a separate tab, a spreadsheet export — measures whether the technology functions, which is not the question. The question is whether it survives the real workflow: interruptions, exceptions, handoffs, the CRM that hasn't been updated since the acquisition. Integration debt discovered in month five was visible in week one to anyone who walked the actual process. - No operator owns it. Pilots sponsored by innovation teams and delivered to no one die the day the innovation team's attention moves on. The owner must be the person whose team lives in the workflow — the one whose numbers improve if it works and who will be embarrassed if it is quietly abandoned. - Adoption was assumed, not designed. Frontline staff do not distrust AI because they fear technology. They distrust it because they own the error. If the tool is wrong and the customer is angry, the human who accepted the output takes the consequence. A pilot that does not design the trust ramp — visible confidence signals, easy correction, a fallback path that costs nothing to use — will report “low engagement” and conclude, wrongly, that the technology wasn't ready. It is the adoption-posture dimension of the readiness score, arriving as a surprise. - Success was defined as sentiment. “The team loves it” is the most dangerous sentence in enterprise AI. Sentiment is real but unbankable. Pilots need one primary metric that moves a business number someone already reports upward — handle time, first-pass yield, days outstanding — and the discipline to ignore applause that arrives without it. A pilot without a kill criterion is not an experiment. It is a subscription — one you pay in credibility as well as budget. ## The 90-day pattern that survives The structure I run — refined across SaaS, healthcare, and automotive-retail engagements — compresses to four phases with hard gates: The 90-day pilot pattern: four phases, each with a gate and the pre-condition it answers. PhaseWorkGatePre-condition it answers Weeks 0–2Baseline the human process; write the definition of doneFour criteria signed by the operating ownerQuality bar, baseline, ownership Weeks 3–6Narrow wedge live in the real workflow, human checkpoint on every outputIntegration debt surfaced while cheapPilot inside the workflow Weeks 7–10Track quality, fully loaded cost per completed task, human acceptance rateThree curves trending, weeklyDesigned adoption, real metric Weeks 11–12Operating review against the pre-committed criteriaScale, kill, or one named extensionSuccess defined in advance ### Weeks 0–2: Baseline and definition of done Measure the human process as it actually runs: cycle time, error rate, exception rate, cost per unit. Write the contract before building: the quality bar, the primary business metric, the kill criterion, the scale criterion. All four signed by the operating owner — not the innovation sponsor, not the vendor. If you cannot get signatures in two weeks, that is the pilot's result. You saved a quarter. ### Weeks 3–6: The narrow wedge, in production Pick the thinnest slice of the real workflow — one document type, one ticket category, one region — and run it live with a human checkpoint on every output. Narrow is the point: a wedge that handles 8% of volume inside the real workflow teaches more than a sandbox that handles 80% beside it. This is where integration debt surfaces while it is still cheap. A wedge is only worth running if it closes a whole loop rather than a fragment. ### Weeks 7–10: Instrument quality, cost, and trust Three curves, tracked weekly: quality against the bar, fully loaded cost per completed task, and human acceptance rate — how often the checkpoint accepts the output unchanged. The third curve is the one teams skip and the one that predicts scale. Quality that isn't trusted delivers nothing. ### Weeks 11–12: The forced decision The gate that defines the whole pattern: on the pre-committed date, against the pre-committed criteria, the pilot scales, kills, or — rarely, and only with a named reason — extends once. The decision is made in an operating review, on the numbers, by the owner who signed in week two. No new evidence is admitted that wasn't defined at the start; that rule exists because purgatory is built from mid-pilot goalpost moves. ## Killing pilots is a capability The organizations that get durable value from AI are not the ones whose pilots always succeed. They are the ones whose pilots always conclude. A clean kill in week twelve — with a baseline, a documented gap, and a written reason — is an asset: it prices the capability curve for your context, and it tells you exactly what has to become true before you try again. It is the same discipline that keeps what compounds and raises an organization's learning rate. An undead pilot teaches nothing and salts the ground for the next attempt, because the second pitch for AI in a workflow is made to an audience that remembers the first. Demos are cheap now. Decisions are the scarce resource. Structure your pilots so a decision is the guaranteed output, and the technology results will take care of themselves. ### FAQ **Why do enterprise AI pilots fail?** Most enterprise AI pilots do not fail; they stall. The causes are product-management decisions made before any code: no pre-committed quality bar, no baseline of the human process, a pilot running beside the workflow rather than inside it, no operating owner, adoption assumed rather than designed, and success defined as sentiment rather than a business metric. **How long should an AI pilot run?** Ninety days, in four gated phases: two weeks to baseline the human process and sign the definition of done, four weeks running a narrow wedge live inside the real workflow, four weeks instrumenting quality, cost per completed task and human acceptance rate, and a final gate where the pilot scales, kills, or extends once with a named reason. **What is a kill criterion for an AI pilot?** A kill criterion is a threshold written and signed before the build that says which result ends the pilot: a quality bar the system must clear, a cost per completed task it must beat, or an acceptance rate it must reach. It is evaluated on a pre-committed date, and no evidence defined after the start is admitted. --- # The Four Jobs of a SaaS CPO in the AI Era URL: https://kevinowens.com/writing/four-jobs-of-saas-cpo-ai-era Published: 2026-05-19 (updated 2026-09-16) Tags: CPO leadership, SaaS leadership, AI strategy Summary: The CPO role is being repriced. Roadmap stewardship and feature-factory management are losing value; four jobs are gaining it — portfolio allocator, quality owner, margin steward, and trust architect. TL;DR - The CPO role is being repriced: roadmap stewardship and delivery choreography are losing value because shipping is no longer the constraint. - Four jobs absorb that value — portfolio allocator, quality owner, margin steward, and trust architect. - Each job is a judgment rather than a process, which is why none of the four can be delegated downward. - Every job has a scorecard question: the last bet killed on its criteria, today's eval score, cost per completed task, and designed behavior at the edge of competence. - Sitting CPOs can audit themselves by counting last week's calendar hours spent on the four jobs versus on choreographing delivery. Having held the Chief Product Officer job four times, I can report that the role has always been unstable — a different shape at every company, renegotiated with every CEO. But what is happening now is not the usual instability. The market is repricing the job itself. Parts of the role that justified the title for a decade are losing value fast, and four jobs are absorbing the value they lose. The parts losing value first: roadmap stewardship and delivery choreography. When AI-assisted teams can ship in days what used to take quarters, being the executive who sequences the backlog and reports the dates is not leadership — it is administration, and administration is exactly what this technology automates. The CPOs who defined their value by owning the process of shipping are discovering that shipping is no longer the constraint. Here is where the constraint moved. Four jobs absorb the value: - Portfolio allocator — deciding how capacity is spread across core, adjacent, and frontier bets, and holding the kill discipline when a bet misses its criteria. - Quality owner — setting the bar each AI behavior must clear and owning the eval culture that enforces it. - Margin steward — carrying AI unit economics as a first-class product metric and designing the packaging ladder with the CFO. - Trust architect — governing the shared, slowly refilled credit line customers extend to the whole product. ## Job 1: Portfolio allocator AI turns product investment from deterministic scheduling into probabilistic allocation. Some fraction of your AI bets will not clear their quality bar regardless of effort, and the capability curve underneath you moves quarterly. Someone has to decide how much of the company's capacity rides on hardening the core, how much on adjacent automation, how much on frontier bets — and, harder, someone has to hold the kill discipline when a charismatic bet misses its criteria. This is capital-allocator thinking applied to product capacity, and it cannot be delegated downward, because every team is structurally in love with its own bet. The AI-era CPO runs the portfolio review the way a good fund runs one: pre-committed criteria, explicit re-pricing, no sentimental positions. It is the same discipline that separates a portfolio from a backlog in the shift from roadmap to bet portfolio, and the same kill criterion that keeps pilots from drifting into purgatory — killing pilots is a capability, not a failure. The scorecard question: can you name the last AI bet you killed on its criteria, on its date? If the answer is none, you are not allocating. You are accumulating. ## Job 2: Quality owner In deterministic software, quality was delegable — QA owned the test suite, engineering owned the bugs. Probabilistic products dissolve that arrangement. The central quality question of an AI feature — how good is good enough for this workflow, and which failures are disqualifying? — is not a testing question. It is a judgment about customers, brand, and risk appetite. It belongs to product, and at the level where it sets precedent, it belongs to the CPO. Concretely, this means the CPO owns the eval culture: every AI behavior ships against a graded eval; the eval encodes product judgment about tolerable versus embarrassing versus disqualifying failures; and eval scores are reviewed in the same forum as revenue, not buried in an engineering dashboard. That is the same artifact the operating model puts at the center of the shift from specs to evals. The companies that treat evals as an ML-team artifact ship quality by accident. The scorecard question: for your most visible AI feature, do you know its eval score today — and did you set the bar it has to clear? The feature-factory CPO managed the process of shipping. The AI-era CPO manages the judgment of what's good, what's economic, and what's trustworthy. Process was delegable. Judgment is not. ## Job 3: Margin steward For twenty years, product leaders could ship value and let finance find the price, because marginal cost was approximately zero. AI ends the free ride: every feature now has a cost of goods that scales with usage, and product decisions — model choice, retry logic, context size, agent autonomy — are margin decisions. The CFO can see the AI bill. Only product can see why it is shaped the way it is and which product choices would bend it. The AI-era CPO therefore carries unit economics as a first-class product metric: cost per completed task per feature, gross margin by packaging rung, cost trajectory as usage mix shifts toward heavier tasks. That metric is the same one that holds a packaging ladder together — see the margin discipline behind the ladder. This job also makes the CPO the CFO's structural partner on pricing — not consulted after, but designing the packaging ladder as part of the product architecture. The scorecard question: can you state the fully loaded cost per completed task of your top three AI features, at the 90th-percentile account? ## Job 4: Trust architect Every AI feature spends customer trust before it earns any back. Customers are extending a provisional line of credit — with their data, their workflows, their tolerance for confident errors — and that credit line is finite, shared across your whole product, and refilled slowly. A single trust incident in one feature raises the adoption tax on every other. Someone has to govern that shared resource: what data the models see and remember, how the product behaves at the edge of its competence, whether failure modes are designed with the same care as success paths, how correction and recourse work when the system is wrong. These decisions cross feature teams, which is exactly why they rise to the CPO. Trust is the retention moat of AI-era SaaS — capability gaps close in months, but a reputation for being safe to rely on compounds for years. The scorecard question: does your product have designed behavior at the edge of its competence, or does it improvise? ## What this means for the people in the job The uncomfortable summary: the CPO job is shifting from managing the production of software to holding four judgments — where to bet, what good means, what it may cost, and what trust requires. Each was always latent in the role. AI makes them the role. The four jobs of an AI-era SaaS CPO: what each replaces, the judgment it holds, and the question that tests it. JobReplacesJudgment heldScorecard question Portfolio allocatorRoadmap stewardshipWhere to betWhat was the last AI bet you killed on its criteria, on its date? Quality ownerDelegated QA sign-offWhat good meansDo you know your most visible AI feature's eval score today — and did you set its bar? Margin stewardPrice set after the factWhat it may costWhat is the fully loaded cost per completed task of your top three AI features? Trust architectPer-feature risk reviewWhat trust requiresIs behavior at the edge of competence designed, or improvised? For sitting CPOs, the audit is simple and bracing: look at last week's calendar and count the hours spent on the four jobs versus the hours spent choreographing delivery. For CEOs hiring product leaders, the interview changes the same way. Ask candidates to walk: - A bet they killed — on pre-committed criteria, on the date, over the team's objection. - A quality bar they set and defended — including the failure they ruled disqualifying. - A margin problem they engineered away — a product choice that bent the cost curve. - A trust decision that cost them a feature — and what it protected. The candidates with crisp answers to those four are the ones built for what the job is becoming — the rest are applying for a role that is being automated out from under them. ### FAQ **What are the four jobs of a CPO in the AI era?** Portfolio allocator, quality owner, margin steward, and trust architect. Each replaces a process the CPO used to run with a judgment only the CPO can hold: where to bet, what good means, what it may cost, and what trust requires. **Why can't a CPO delegate AI quality to the ML or QA team?** Because the central question — how good is good enough for this workflow, and which failures are disqualifying — is a judgment about customers, brand, and risk appetite, not a testing question. The CPO owns the eval culture and the bar each eval has to clear. **Should the CPO own AI unit economics, or the CFO?** Both, from different ends. The CFO can see the AI bill; only product can see why it is shaped that way and which product choices — model, retries, context size, agent autonomy — would bend it. That makes cost per completed task a product metric and the CPO the CFO's partner on packaging. --- # The AI Readiness Score: Assess a Product Org in 2 Weeks URL: https://kevinowens.com/writing/ai-readiness-score Published: 2026-04-28 (updated 2026-09-16) Tags: AI strategy, Diagnostics, Operating models Summary: Before recommending any AI investment, I score the organization on five dimensions. The score predicts outcomes better than the technology choice ever does — and it tells you your first move. TL;DR - Readiness, not use-case selection, is the first decision: the same use case compounds in a ready organization and becomes an undead pilot in an unready one. - Five dimensions — data foundations, workflow clarity, decision rights, capability density, adoption and trust posture — are each scored 1 to 5 from evidence over two weeks. - The useful output is not the composite number but the lowest dimension, which names your first project — and it is almost never a model. - Composite below 2.5 means build foundations; 2.5 to 3.5 means one narrow wedge run properly; above 3.5 means the constraint is ambition, not readiness. - Scores are read from artifacts and a workflow walk rather than self-report, because self-assessment inflates every dimension by roughly a point. The first question leadership teams ask me is almost always some version of “which AI use case should we start with?” It is a reasonable question, and it is premature. Use-case selection is the second decision. The first is an honest reading of what the organization can currently metabolize — because the same use case that compounds inside a ready organization becomes an undead pilot inside an unready one. So every engagement starts the same way: two weeks, five dimensions, each scored 1 to 5. The composite — I call it the AI Readiness Score — predicts outcomes better than any technology choice, and more usefully, the lowest dimension tells you exactly what to do first. (A three-question quick version lives in the Lab; this essay is the full instrument behind it.) ## The five dimensions - Data foundations. Not “do you have a data lake” — most companies scored highly by their own data teams still fail this dimension where AI actually needs it. The question is whether the data that describes your core workflows is accessible, accurate, and fresh enough to act on. A 2 looks like: the answer to any operational question requires an analyst and a week. A 4 looks like: the systems of record agree with each other, someone owns each critical dataset, and a new integration is a task, not a project. You cannot retrieve your way around records that are wrong. - Workflow clarity. AI automates workflows, and you cannot automate what you cannot describe. A 2 is an organization where the process lives in the heads of three tenured operators and every walkthrough produces a different diagram. A 4 is an organization that knows its volumes, exception rates, and handoffs — where someone can say “step four is 30% of cycle time and 70% of errors” with numbers behind it. Low workflow clarity is the most common silent killer: the pilot gets built against the imagined process and dies against the real one, which is the first pre-condition of pilot purgatory. - Decision rights and governance. When an AI system produces a wrong answer with consequences, who finds out, who decides what changes, and how fast? A 2 is governance by escalation and improvisation — every incident convenes an ad-hoc meeting of everyone. A 4 has boring, written answers: quality bars set in advance, an owner per deployed system, a defined path from “the model is wrong” to “the model is fixed.” Teams over-invest in AI ethics theater and under-invest in this mundane operational version, and it is the mundane version that determines whether anything reaches production. - Capability density. Not headcount with AI titles — fluency in the teams that own the work. Can the product trio closest to your core workflow read an eval, reason about failure modes, and estimate cost per task without filing a ticket to a specialist group? A 2 is an organization where all AI knowledge sits in one team that everything queues behind. A 4 is one where the median product team has shipped and operated something probabilistic, however small. Density beats specialization because compounding happens where the work is, not where the expertise is warehoused. - Adoption and trust posture. The dimension every technical assessment skips, and the one that most often decides the outcome. What happened to the last three tools rolled out to the frontline — adopted, tolerated, or quietly worked around? Do operators own the errors of tools they are told to use? Is there a correction path that costs nothing socially? A 2 is a workforce with scar tissue from transformation programs past; a 4 is one where frontline teams have visibly shaped a rollout before and trust that a fallback path exists. AI lands on whatever trust surface already exists. It does not create one. The AI Readiness Score: five dimensions, each scored 1 to 5 DimensionA 2 looks likeA 4 looks like 1. Data foundations Any operational question needs an analyst and a week. Systems of record agree, each critical dataset has an owner, a new integration is a task. 2. Workflow clarity The process lives in three tenured heads; every walkthrough draws a different diagram. Volumes, exception rates, and handoffs are known with numbers behind them. 3. Decision rights and governance Governance by escalation; every incident convenes an ad-hoc meeting of everyone. Quality bars set in advance, an owner per deployed system, a defined path to a fix. 4. Capability density All AI knowledge sits in one team that everything queues behind. The median product team has shipped and operated something probabilistic. 5. Adoption and trust posture A workforce carrying scar tissue from transformation programs past. Frontline teams have shaped a rollout before and trust that a fallback path exists. The score's real output is not the number. It's the sequence: your lowest dimension is your first project — and it is almost never a model. ## How the two weeks run The method is deliberately unglamorous. Week one: read the artifacts the organization already produces — roadmaps, incident channels, ops dashboards, the last three pilot post-mortems if they exist (their absence is itself a data point) — and walk the top two workflows end to end, sitting with the people who run them. Week two: structured interviews across the leadership seam — product, engineering, ops, finance, frontline management — scoring each dimension from evidence, not self-report. Self-assessment inflates every dimension by roughly a point; the artifacts and the workflow walk are the correction. ## Reading the composite - Below 2.5 — build foundations, don't buy capability. An organization here that funds an ambitious AI program is pre-paying for a purgatory pilot. The right spend is on the lowest one or two dimensions: instrument the workflow, fix the record systems, name the owners. Six months of unglamorous work moves the score a full point and changes what every later dollar buys. - 2.5 to 3.5 — one narrow wedge, run properly. Ready enough to learn from production, not ready to parallelize. Pick a single high-clarity workflow, run the disciplined 90-day pattern with a hard scale-or-kill gate, and use the pilot deliberately as a readiness-building instrument — it should raise governance and adoption scores whether or not it scales. - Above 3.5 — the constraint is ambition, not readiness. Organizations here usually under-invest out of habit, running cautious pilots when they could be re-architecting workflows. The work at this level is portfolio construction and pace: multiple concurrent bets, explicit horizons, and kill discipline. A pattern worth naming: scores cluster tighter than teams expect. The modal mid-market organization I assess lands between 2.0 and 2.5 — genuinely capable people, real data assets, and one or two foundational gaps that would silently cap any AI investment made on top of them. Telling a leadership team eager for an AI roadmap that their first project is workflow instrumentation is not the fun version of this job. It is the useful version. The score exists to make that conversation a matter of evidence instead of opinion — and to make the next assessment, a year later, show a number that moved. ### FAQ **What is an AI readiness assessment?** It is a structured read of whether an organization can absorb an AI investment, scored before any use case is chosen. This one takes two weeks and scores five dimensions — data foundations, workflow clarity, decision rights, capability density, and adoption posture — from artifacts and a workflow walk rather than self-report. **How do you know if your company is ready for AI?** Score the five dimensions 1 to 5 and read the composite. Below 2.5 you are not ready to buy capability and should fix foundations; 2.5 to 3.5 you are ready to learn from one narrow wedge in production, not to parallelize; above 3.5 readiness is no longer the constraint and ambition is. **What should our first AI project be?** Your lowest-scoring dimension, not your most attractive use case. Low workflow clarity means instrumenting the workflow; weak data foundations means fixing the systems of record; thin governance means naming owners and quality bars. The two-week read exists to make that assignment a matter of evidence instead of opinion. --- # Ship the Loop, Not the Fragment URL: https://kevinowens.com/writing/ship-the-loop Published: 2026-04-07 (updated 2026-09-16) Tags: Execution, Product craft, Operating models Summary: The operating discipline behind everything I build: ship the smallest complete loop from input to value, measure what it changes, keep what compounds — and kill what doesn't. TL;DR - A fragment is any work that needs imagination to connect it to user value; a loop is the smallest complete, instrumented path from a real input to delivered value. - For AI products a loop is only complete when it carries four segments: an eval, a fallback, feedback capture, and a cost meter. - Instrument what the loop changes — one pre-chosen business number someone already reported — not what it does. - Compounding loops get doubled down on; everything else gets a dated, written kill with the re-open conditions recorded. - One weekly hour per loop owner — four numbers, three questions — turns the portfolio into an allocation decision instead of a status report. Every operating principle I hold reduces to three sentences: ship the loop, not the fragment. Measure what it changes, not what it demos. Keep what compounds — and kill what doesn't. This essay is the long version, because the short version gets nodded at and then ignored, and the gap between nodding and doing is where product organizations go to underperform. ## Fragments versus loops A fragment is any piece of work that requires imagination to connect to user value: the beautiful interface with mocked data, the model that performs in the notebook, the integration that works in staging, the strategy deck with no owner for slide twelve. Fragments are seductive because they demo well, and demos are the currency of internal status. An organization can stay busy for years shipping fragments — each one praised, none of them compounding — while carrying the warm feeling of momentum. A loop is the smallest complete path from a real input to delivered value, instrumented well enough to tell you the truth. Emphasis on every word: - Smallest — because scope is the enemy of completion. - Complete — because a loop with one missing segment delivers exactly nothing; value is not proportional to percent-done. - Instrumented — because an unmeasured loop is just a fragment with better marketing. The discipline sounds obvious. It is not practiced, because loops force you to eat the unglamorous parts first — the data plumbing, the edge cases, the handoff nobody wants to own — and fragments let you defer them indefinitely. Show me a team's last quarter and I can tell you which diet they are on: loops produce small, complete, slightly boring wins that stack; fragments produce impressive reviews and a strange absence of changed numbers. ## What a loop means for AI products AI raised the stakes on this discipline, because AI is the greatest fragment-generating technology ever built. The distance between a stunning demo and a dependable system has never been wider, and the demo has never been cheaper. For an AI feature, the complete loop includes four segments that fragment-thinking omits: - The eval. A graded definition of good, runnable on every change. Without it you cannot tell whether the loop is improving or drifting — you can only tell whether people still clap. This is the same artifact that replaces the spec in the shift from specs to evals, and the one a CPO owns as quality owner. - The fallback. Designed behavior at the edge of competence: how the system degrades, hands off, or declines. A loop that only works when the model is right is a fragment wearing a loop's clothing. - The feedback capture. Corrections, rejections, and escapes harvested from inside the workflow, feeding the next quality target. This is the segment that makes the loop an asset instead of an artifact. - The cost meter. Fully loaded cost per completed task, visible to the team that owns the loop. Unit economics is part of the loop's truth, not finance's problem — the same number that carries the margin discipline behind a pricing ladder. The four segments of a complete AI loop SegmentWhat it isWhat is missing without it EvalA graded definition of good, runnable on every changeAny way to tell improvement from drift FallbackDesigned behavior at the edge of competenceDependability when the model is wrong Feedback captureCorrections and escapes harvested inside the workflowCompounding — the loop stays an artifact Cost meterFully loaded cost per completed task, owned by the teamUnit economics, and the right to price the loop A loop that runs is worth more than a platform that impresses. You can compound a running loop. You can only present a platform. ## Measure what it changes The second sentence exists because instrumentation is routinely aimed at the wrong target. Teams measure what the loop does — requests served, outputs generated, sessions engaged — because those numbers are easy and always go up. The question that matters is what the loop changes: a business number someone already reported before your feature existed, moving in the direction you predicted, by an amount you would defend in front of the CFO. Handle time. First-pass yield. Conversion. Days outstanding. One primary metric per loop. Chosen before launch, not discovered after — a metric selected retrospectively is an alibi, not a measurement. And the corollary that keeps teams honest: if no business number could plausibly move, the loop should not be built. “Strategic” is not an exemption; it is the word fragments use to apply for permanent funding. ## Keep what compounds — and kill what doesn't The third sentence is portfolio discipline applied weekly. Compounding loops share a signature: each cycle of usage makes the next cycle better — more feedback, better evals, lower cost, higher trust — without proportional new investment. These deserve more than maintenance; they deserve doubling down, because compounding assets are rare and the instinct to move on to the next new thing systematically starves them. Everything else is a candidate for the kill list, and the kill needs to be an act, not an absence: a dated decision, a written reason, instrumentation archived, the team redeployed with credit rather than stigma. Quiet abandonment — the fragment's natural death — teaches the organization nothing and leaves a residue of zombie surfaces that all cost trust to maintain. A clean kill teaches the organization what the bar is, which is why killing pilots is a capability rather than a failure. In AI work especially, where the capability curve moves underneath you, a documented kill also leaves behind the one thing a dead project can bequeath: the precise conditions under which the bet becomes worth re-opening. ## The weekly loop review The ritual that operationalizes all three sentences fits in one recurring hour. Each loop owner brings four numbers: - Eval score against the bar. - Primary business metric against baseline. - Cost per completed task, fully loaded. - One week-over-week delta they are trying to move. And answers three questions: - Is it improving? Eval and business metric, both against a number from before. - Is it compounding? Does each cycle of usage make the next one better without proportional new investment? - What would make us stop? The kill criterion, written down while everyone is calm. The review's output is allocation — which loops get more, which get maintenance, which get a kill date. That hour replaces most of what product organizations currently do in status meetings, because status describes activity and the loop review prices assets. Run it for a quarter and the portfolio sorts itself: the fragments become embarrassing to present, the compounding loops become obvious to fund, and “learning rate” stops being a phrase from an essay and becomes a number your team knows it is managed on. Ship the loop. Measure the work. Keep what compounds. Everything else in my practice is commentary on those three sentences. ### FAQ **What is the difference between shipping a loop and shipping a feature?** A feature can be a fragment: a surface that demos well but still needs imagination to connect it to value. A loop is the smallest complete path from a real input to delivered value, instrumented well enough to tell you the truth. If one segment is missing, the loop delivers nothing — value is not proportional to percent-done. **What does a complete loop include for an AI feature?** Four segments fragment-thinking omits: an eval, a graded definition of good runnable on every change; a fallback, designed behavior at the edge of competence; feedback capture, corrections harvested inside the workflow; and a cost meter, fully loaded cost per completed task visible to the team that owns the loop. **How do you decide when to kill an AI project?** Kill anything that is not compounding, and make the kill an act rather than an absence: a dated decision, a written reason, instrumentation archived, the team redeployed with credit rather than stigma. Record the conditions under which the bet becomes worth re-opening. Quiet abandonment teaches the organization nothing and leaves zombie surfaces that cost trust to maintain.