Your AI Programme Didn't Fail. Your Governance Did.

14 April 2026 · Prashant Chamarty

A CDO I spoke with recently put it this way: “The technology worked. The business case didn’t. We still got cancelled.”She is not alone. And the reason is rarely what the post-mortem says it as it is. Post-mortems blame data quality, change management, or insufficient executive sponsorship. These are real frictions. They are not the primary failure. The primary failure happens before a single line of code is deployed in the room where the investment is approved.

When an AI investment is approved, three legitimacy frameworks operate simultaneously and none of them are aligned.

The C-suite approves the use of strategic logic: competitive positioning, market differentiation, and future optionality. Long horizon. Qualitative. Inherently hard to falsify. The finance function evaluates using the framework it already manages: cost per unit, headcount productivity, and EBITDA contribution. Short horizon. Quantitative. Designed to evaluate efficiency investments. The deployment team measures using technical proxies: model accuracy, adoption rate, and the number of tasks automated. Local. Operational. Disconnected from business outcomes by design.

These frameworks are not just different. They are structurally incompatible. A programme that succeeds by C-suite logic: we are building a data capability our competitors cannot replicate in five years; will routinely fail by finance logic; cost per transaction has not moved in twelve months; and will get cancelled. The technology performed. The agreement failed. That is not a measurement problem. It is a governance problem. No amount of better dashboards solves a governance problem.

What you need to agree on first

Before any measurement infrastructure, attribution model, or ROI framework, one question must be answered, and answered the same way by every stakeholder who will judge the investment:

What kind of value are we claiming?

There are five fundamentally distinct answers. I call them Value Logics. Each requires a different proof, a different time horizon, and a different governance owner. They cannot be averaged or traded off against each other at renewal without producing a verdict that means nothing.

VL1 - Efficiency Logic

We do the same work with fewer resources.

The most common AI value claim. Also, the most likely to produce a real gain that never reaches the P&L. Here is the mechanism. AI automates a task. The task sits inside a process designed for human throughput. The process does not change. The labour cost does not change. The productivity gain stays trapped at the task level absorbed as capacity slack, converted to more work by the same team, or relabelled as “higher quality output” with no cost reduction attached. The fundamental truth: automating a step within an unchanged process does not reduce the process’s cost. It reduces the time spent on one step, which the system immediately fills with something else. Until the process is redesigned around the new capability, the savings exist only on a slide. Klarna achieved genuine VL1 efficiency, AI displacing agent cost at scale. Then the company began rehiring human agents as quality degraded on complex queries. Even the clearest efficiency case has a sequel. Task automation without process redesign creates a downstream bottleneck that eventually costs more than the upstream savings delivered.

VL1 only works when process redesign is a precondition for deployment, not a follow-on phase. If the process is not redesigned before the AI is switched on, you are not claiming efficiency logic. You are claiming that task automation will somehow become P&L savings without anyone designing that pathway. That is not a business case.

The metric: cost per unit of output at the P&L line, not hours saved, tasks completed, or satisfaction scores.

Who governs it: the CFO, with agreed P&L line ownership confirmed before deployment begins.

VL2 - Augmentation Logic

Same people. Better judgments.

This is where most knowledge-worker AI lives: legal, risk, compliance, strategy, and research. The premise is not headcount reduction. We expect the quality of decisions made by the same people to improve and produce measurable outcomes over time. The fundamental problem is epistemological. You cannot prove that a decision made with AI assistance was better than the same decision made without it. You are measuring a counterfactual. Counterfactuals do not appear in the finance function’s KPI framework. This creates a predictable failure sequence. The organisation deploys an augmentation tool. At renewal, finance evaluates using KPIs it already manages: cost per unit, headcount, and task throughput. These items don’t move; the tool isn’t designed to move them. Finance concludes the investment underperformed. The programme is cut.

Microsoft’s M365 Copilot is the most visible live case. After two years across the world’s largest enterprise productivity platform, adoption remains a fraction of licensed users. Microsoft responded not by improving the product, but by cutting the price and changing the commercial model. The product works. Because the value logic wasn’t agreed upon, it cannot be proven. Finance evaluated it with efficiency tools. Efficiency tools found no efficiency. The solution is a pre-specified decision-quality framework used before purchase. What decisions does this function make regularly? How do we assess whether they are good? What would measurably better look like in eighteen months? If those questions cannot be answered before the tool is bought, VL2 has not been defined. An undefined value cannot be defended.

The metric: decision quality score, defined, baselined, and owned before deployment.

Who governs it: the business unit leader, not finance. Assigning VL2 evaluation to finance is the governance error, not a finance failure.

VL3 - Advantage Logic

We can do what our competitors structurally cannot.

This is the value logic most commonly approved using C-suite reasoning and most commonly handed to finance at renewal. That handover is where programmes die. IBM Watson Health is the proof. Hospitals purchased Watson for a clinical outcome advantage. The measurement framework applied at renewal was processing volume and NLP accuracy. When MD Anderson cancelled, the technology had done exactly what it was designed to do. Nobody had written down at investment approval: this programme succeeds when clinical outcomes improve by a defined measure, evaluated by a named function, at a specified date. Without that agreement, finance applied what it had. What it had was wrong.

The competitive advantage case is genuine when two conditions hold. First: the data that generates the advantage is proprietary, and competitors cannot replicate it by starting today. BlackRock’s Aladdin runs on data generated by managing assets at a scale that took decades to accumulate. Rolls-Royce’s predictive maintenance relies on engine operational data that accrues only from flying those engines; no competitor can collect it retroactively. The advantage is not the AI model. It is the data the model runs on, and the time required to generate it. Second: the organisation is actively compounding that advantage rather than maintaining a static position. The critical challenge: AI itself does not confer a competitive advantage when every competitor has access to the same foundation models at the same cost. What confers advantage is what the AI runs on: proprietary data, process knowledge, and customer relationships that cannot be replicated. If the VL3 claim depends on the AI being uniquely powerful rather than the underlying asset being uniquely theirs, it will not survive a rigorous board challenge.

The metric: competitive position indicators, win rate, switching cost created, and capability gap versus named competitors. Not cost per transaction.

Who governs it: the board, with a regular strategic review cadence. Placing VL3 governance with finance guarantees it will be evaluated with the wrong tools.

VL4 - New Value Logic

Revenue lines that did not exist before.

Some AI investments create genuinely new products, services, or business models. This requires the most honest business case and consistently produces the most optimistic one.The omitted question, almost every time: what happens to the revenue of the offering this new offering competes with? Chegg built a substantial business helping students with coursework. When students adopted AI assistants that did the same job at no cost, Chegg’s revenue collapsed, not because Chegg failed to innovate, but because the value customers were paying for became freely available without them. An honest business case should have included a model of the cannibalisation risk. It rarely did.

The same question faces any enterprise VL4 investment. An insurer building AI-native parametric products must ask whether those products replace existing policies. A bank building an AI-powered financial planning must ask whether customers consolidate relationships in ways that reduce revenue from other products. A manufacturer building predictive maintenance as a service must ask whether it cannibalises the parts and labour revenue that the programme was designed to protect. VL4 demands a business case with an explicit cannibalisation model. What existing revenue is at risk? Net of substitution, what does the value case look like at a defined point in time? A VL4 business case without this is not complete.

The metric: new revenue attributable to AI-enabled products, net of cannibalised existing revenue and margin differential.

Who governs it: the CEO and board, with cannibalisation modelling at investment approval,l not the product function alone, which carries a structural incentive toward optimistic projections.

VL5 - Insurance Logic

We can withstand disruptions that our competitors cannot.

This is the only value logic that rarely yields to standard ROI evaluation. It must be evaluated using scenario analysis, and it is the one most commonly approved without either. The mechanism: an organisation builds AI capability before a disruption arrives. When it comes to regulatory, competitive, and operational matters, a prepared organisation absorbs them. The unprepared competitor does not. The investment’s value is not what it returns during normal operations. It is the asymmetry it creates when conditions change.

The organisations that navigated COVID’s supply chain disruption with the least damage were those that had built AI-powered demand forecasting before 20,20, not because they predicted a pandemic, but because they had built resilient systems as part of an operational discipline. When disruption arrived, those systems adapted. Traditionally structured operations did not. The European regulatory environment makes VL5 concrete right now. DORA is in force. The EU AI Act’s high-risk requirements arrive through 2027. Organisations that built AI governance infrastructure before these mandates took effect are absorbing compliance into routine work. Those that didn’t are building the same infrastructure under enforcement pressure, at emergency cost. The value of having built it early cannot be expressed as ROI on a specific investment; rather, it is the return on a governance decision made under uncertainty. Finance evaluates investments against defined returns on defined time horizons. VL5 investments return value contingent on events that may not occur within any given measurement period. Asking finance to approve VL5 is asking the wrong function. The board sets the organisation’s risk appetite. VL5 is a board-level investment in that risk posture.

The metric: scenario resilience is the organisation’s tested capacity to maintain operations across defined disruption scenarios.

Who governs it: the board. It sets the scenarios, approves the investment, and reviews against scenario outcomes, not quarterly P&L movements.

The Five Disciplines

The value logics are diagnostic. These are operational issues that must change before investment is approved.

D1 - Name one logic. Hold to it.

The sponsoring executive names the primary value logic and defends it against the alternatives before approval is granted. If the logic isn’t clear at the time, name a primary and an exploratory secondary, with different success criteria and time horizons for each. What is indefensible is naming none. Organisations that claim all five logics simultaneously are describing possibility rather than committing to a thesis. An investment without a committed logic will be evaluated by whatever framework the reviewer happens to apply. That framework will almost certainly be wrong.

D2 - Measurement infrastructure is a precondition for deployment, not a project phase.

Measurement infrastructure is almost universally specified after the vendor contract is signed. By this point, the deployment team has no incentive to delay, and the pre-deployment baseline has been permanently lost. Without a baseline, there is no before. Without a before, there is no proof of after. The governance fix is categorical: no baseline established, no deployment approved.

D3 - Design attribution before the AI is switched on.

AI value compounds with process change, data improvements, and workforce behaviour shifts simultaneously. When a risk AI reduces credit losses, some reduction comes from the model, some from process changes, and some from how analysts changed their work. If the attribution framework is not agreed upon before deployment, the causal chain cannot be reconstructed afterwards. The counterfactual is gone.

D4 - Report by logic. Prohibit mixed programme reviews.

When efficiency metrics, decision-quality scores, and competitive indicators appear in the same board pack, the result is noise that appears as underperformance across all logics. Finance reviews VL1. Business unit leadership reviews VL2. The board reviews VL3 and VL5. Mixed reviews produce verdicts that are simultaneously unfair to the programme and useless to the board.

D5 - Set your own review date before deployment begins.

The renewal cliff arises because organisations let the finance review cycle designed for operational expenditure, not for capability investments that compound over years, define when AI is evaluated. The antidote: the review date, success criteria, responsible function, and required evidence are documented at the time of investment approval. Boards that set their own accountability moments are the ones that navigate them.

The decision that actually matters

Most AI governance conversations focus on model selection, data quality, and deployment tooling. These are real challenges. They are not the primary ones.

The primary challenge is that the people who approve AI investments, the people who deploy them, and the people who evaluate them at renewal operate from fundamentally different frameworks and organisations rarely address this before committing the budget.

One honest limit: five value logics create five potential escape routes from accountability. The antidote is D1, applied with disciplined use of one primary logic: name it before investing, and hold it at renewal, even if there’s pressure to reclassify.

The leaders making genuine progress on AI are not those with the best models or the most sophisticated measurement tools. They are the ones who first agreed on what they were claiming, how they would prove it, who would own the verdict, and when it would be rendered.

That is a governance capability. Not a technology capability. And it is almost absent from how enterprise AI investments are currently governed.

Before your last AI investment was approved, did everyone who would judge it at renewal agree on which logic it was claiming, and who owns that verdict?

If the answer is no, you already know where the risk sits.

Where does this hold in your experience, and where does it break? I’d particularly like to hear from those who’ve been through a renewal cycle and seen the misalignment first-hand.


Originally published on LinkedIn.

← Back to all essays