What Karpathy Didn't Say About Your Enterprise AI Strategy - Part 1: The Diagnostic

12 May 2026 · Prashant Chamarty

Your enterprise AI rollout is probably sequenced wrong. Not because of the models. Because of the operating model.

MIT NANDA estimates 95% of enterprise GenAI pilots are failing to deliver meaningful ROI. The capital is moving. The value is not. I have spent years operating at the intersection of hyperscalers, GSIs, and the regulated enterprises, trying to deploy AI. What I keep seeing is not a technology problem. We’re applying a deterministic-era operating model to a non-deterministic technology, which is leading to predictable failure.

Andrej Karpathy’s recent talk gave builders a framework. Read against enterprise reality, it reads like a diagnostic. Here are the four failure modes I see most consistently.

01 / YOUR PROCUREMENT TEAM IS BUYING SOFTWARE 1.0 AND CALLING IT AI

Karpathy describes Software 3.0 as the era where the context window is the programming surface. The unit of value has shifted from the application to the intelligence embedded in it. Enterprise procurement has not made the same shift.

RFPs are still scoring seat counts, SLA guarantees, and integration depth. These are the right criteria for Software 1.0. These criteria aren’t appropriate for an AI system where intelligence resides in prompt composition, the learning loop, and domain adaptation, none of which are typically included in a standard vendor scorecard.

The buying motion is structurally mismatched to what is being sold.

02 / WHAT HAPPENS TO YOUR DIGITAL TRANSFORMATION WHEN THE APP SHOULDN’T EXIST?

Karpathy built a complex multi-component app to photograph a restaurant menu and generate food pictures. Then a single Gemini prompt rendered the same output directly. His conclusion: “The app shouldn’t exist.”

The enterprise version of this question: which of your current applications would survive if an agent could query your systems without them?

Any application whose primary function is moving structured data between humans and a database is at risk of becoming unnecessary middleware. The agent reads the database directly. The layer in between becomes optional.

I am not saying your SAP investment is obsolete. The boards approving 2026 renewals are not asking this question, and they should be.

ServiceNow paid $2.85B for Moveworks. Salesforce launched Agentforce inside ChatGPT. Both are defensive moves against exactly this threat. The incumbents see it clearly.

03 / THE 19-POINT PENALTY, DEPLOYING AI INTO THE WRONG CIRCUITS

Karpathy’s “jagged intelligence” describes a capability surface that is not smooth. Models peak in verifiable domains, code, math, and structured data, and are rough everywhere else.

Dell’Acqua and BCG’s research found something that should stop every CDO in their tracks: AI users on tasks outside the capability frontier performed 19 percentage points worse than non-users. Not neutral. Actively worse.

Klarna reversed its AI customer service deployment. McDonald’s shut down its IBM drive-thru AI. Air Canada lost a misrepresentation case over its chatbot. Each was a deployment into a high-variance, low-verifiability workflow. Each was a wrong-circuit deployment.

Enterprises deploy AI by workflow proximity: “This role is language-heavy, so AI fits.” That is not topology mapping. It is a guess.

04 / VERIFIABILITY, THE AI ROLLOUT CRITERION ALMOST NOBODY IS USING

Karpathy’s framework has a direct implication for enterprise sequencing: automate first what you can verify, regardless of how visible or cost-saving it appears.

Over half of enterprise GenAI budgets flow to sales and marketing. The highest measurable ROI is in back-office automation. The IRS reduced the case-opening time from 10 days to 30 minutes. JPMorgan saves 360,000 work-hours annually on credit agreement review. Both are verifiable, structured, high-frequency workflows.

Financial services have the richest portfolio of verifiable AI domains in the economy. So does advanced manufacturing. So does the retail supply chain. None of these leads the board deck. The customer-facing chatbot does.

The cost-savings lens is borrowed from RPA and BPO, when the technology was deterministic. LLMs are stochastic. Applying deterministic-era criteria to stochastic technology produces Klarna outcomes.

THE PATTERN ACROSS ALL FOUR

The procurement mismatch, the obsolescence problem, the wrong-circuit deployment, and the wrong sequencing criterion share a common root: a deterministic-era operating model applied to non-deterministic AI.

The models are not the problem. The operating model is.

Part 2 covers the four structural changes that follow, the quality bar problem, governance for non-enumerable systems, where AI spend actually belongs in the budget, and the proprietary advantage most enterprises are leaving on the table.

Of these four failure modes, which one is your organisation most exposed to right now?

I’m particularly interested in hearing from CDOs and CIOs in Financial Services, Retail, and Manufacturing, where verifiability is highest, and the gap between AI promise and production reality is most visible.

#EnterpriseAI #AIStrategy #CDO #AIGovernance


Originally published on LinkedIn.

← Back to all essays