The Most Dangerous AI Is the One That Does Exactly What We Asked

22 September 2026 · Prashant Chamarty · 16 min read

operating-model agents controls governance decision-ai

TLDR

The biggest problem with enterprise AI is not hallucination. It is obedience. We trained models to please a person in the loop, then began removing the person and calling the result an agent. Real automation serves code, not a chat window: bounded decisions, structured state, inspectable policy, and escalation paths. A probability is not governance. The choice menu is part of the model.

Decision AI begins where the copilot metaphor breaks: when nobody is left to catch the model pleasing us.

The biggest problem with enterprise AI is not hallucination.

It is obedience.

We have spent three years teaching models to be useful, agreeable and fluent in front of a person. Then we began removing the person and calling the result an agent.

That is not a small product upgrade. It changes the meaning of success.

A copilot can be impressive while being wrong because a human is still holding the controls. An autonomous system can be impressive, persuasive and fast while quietly scaling the wrong action. The first failure is annoying. The second becomes an operating model.

This is why “better chat” is the wrong frame for the next phase of enterprise AI. The hard problem is no longer getting a model to answer. It is getting a system to make a bounded decision, expose uncertainty, obey policy and remain accountable when no one is watching every move.

That is the opening for Decision AI.

Not because every enterprise problem is a decision problem. And not because a probability is automatically more truthful than a paragraph. It matters because assistance and automation are different workloads, and we have been pretending they are the same.

The distinction became clearer to me after listening to TypeSafe founder Diogo Almeida describe Jev not merely as a decision model, but as the first of a broader class of “large programmable” or “System One” models. His definition matters: the consumer is code, not a person.

That one shift changes almost everything.

The hidden contract behind every copilot

A copilot has a silent safety mechanism: the user.

The model drafts. The person checks. The model proposes. The person approves. The model gets something subtly wrong, and the person is expected to notice.

This arrangement explains both the magic and the frustration of today’s systems. They are optimised for a world in which usefulness is judged by a human in the loop. Fluency matters. Relevance matters. Confidence often feels helpful. The interaction succeeds when the person says, “Yes, that is what I meant.”

But autonomy changes the contract. There may be no person inspecting the output before it routes the claim, freezes the payment, prioritises the patient, suspends the seller or approves the refund.

Almeida’s argument is more precise than “agents will replace copilots.” Models trained to produce responses that people prefer are optimised for an interaction. Automation is infrastructure. A model that sits inside software, runs in the background and feeds another system cannot rely on charm, clarification or a patient operator who works around its quirks.

A chatbot refusal is inconvenient. A dependency that refuses unpredictably can break the process that called it. A chatbot can ask the person what they meant. Background software may have no person to ask.

This does not mean safety should disappear. It means safety has to move from personality to architecture: explicit permissions, bounded inputs and outputs, policy checks, deterministic controls, audit trails and escalation paths.

The enterprise version is simple:

If a model is rewarded for producing an answer a person likes, do not assume it is ready to produce an action the business can defend.

Real automation is not an assistant without a screen

We use the word “agent” too loosely.

An assistant waits. It receives an instruction, produces something useful and hands control back. Even when it calls tools, its natural unit of work is still a conversation with a person.

Real automation is different. It runs when no one is looking. It can be called thousands or millions of times by code. Its output becomes an input to another system. It must behave like a dependable component rather than a talented colleague having a good day.

This is the frontier Almeida is pointing towards: intelligence that disappears into software.

That is also why the economics change. A person may write a handful of prompts. A useful background process may make the same narrow judgement millions of times, then become a dependency for higher-level workflows. The value comes less from one spectacular answer and more from a small decision that can be repeated, measured and composed.

The right test for an agent is therefore not, “Did it complete the demo?”

It is:

If not, it may be a useful assistant. It is not yet automation.

A decision is not a short generation

Most enterprise AI systems still turn every task into text generation, even when the required output is tiny.

Read this case. Write JSON. Pick one of four routes. Add a confidence score. Try again if the JSON breaks.

That can work, but it confuses the wrapper with the workload.

A generative workload has an open output space: explain this incident, draft a response, summarise the evidence, propose a plan. A decision workload has a bounded action space: allow or hold, route to A or B, approve or escalate, choose a reserve band, ask for one of three missing documents.

The distinction is not absolute. Decisions often need generative work upstream to extract evidence and downstream to explain an outcome. But the final act has a different shape:

  1. inspect a defined state
  2. choose among permitted outcomes
  3. expose uncertainty
  4. apply a policy that prices the consequences
  5. record what happened

The fifth step is the one most demos omit.

A model returning approve: 0.82 has not approved anything. It has produced an estimate. The enterprise still has to decide whether 0.82 is enough, which errors matter, what happens to the other 0.18, and who owns the result.

This is where Jev becomes interesting. Almeida describes a machine-native model whose outputs are meant to be consumed directly by code. Archer Hume’s black-box investigation suggests an architecture designed around shared state, isolated questions, interacting option sets and direct probability distributions rather than token-by-token prose. His strongest observations are behavioural, not architectural: common evidence appears reusable across many questions; options influence one another; and large batches of short decisions can be returned quickly.

The exact internals remain unverified. Hume’s sparse mixture-of-experts hypothesis is the weakest part of the reconstruction. But the useful question is not whether an architecture diagram is literally correct. It is whether the service behaves better for the decision contract we actually need.

There is also a useful correction to the language. “Decision model” may be too narrow. Almeida calls Jev a System One model: fast, machine-native intelligence for software, with decisions as the first visible use case. That broader claim is still a vendor thesis, not a settled category. But it makes the ambition clearer. This is not a better way to print JSON. It is an attempt to give code a native way to use intelligence.

Decomposition is the real engineering move

The most practical idea in Almeida’s talk is not about model architecture. It is about software design.

Break the task down.

Do not ask one model to read the entire state, absorb a page of instructions and somehow produce the final business action. Split the workflow into the smallest meaningful judgements. Ask which facts are present. Ask which risk indicators apply. Ask whether evidence conflicts. Ask whether the case is in scope. Then let code combine those answers under an explicit policy.

This is more than prompt engineering. It changes how failure behaves.

A large prompt creates one opaque surface. When it fails, the team often adds another sentence and hopes the instruction is remembered next time. A decomposed workflow creates smaller units that can be evaluated separately. When a new failure appears, the team can add a question, a test case, a threshold or a deterministic rule.

That is much closer to software engineering.

It also makes evaluation possible. “Did the agent handle the case well?” is hard to measure. “Did it detect the missing invoice?” or “Did it classify this request as outside the refund window?” is much easier.

The design principle is simple:

A background agent should be built from decisions small enough to test, not instructions large enough to impress.

Almeida also makes a sharp point about state. We keep converting structured information into giant strings because language models began as conversational tools. But if code is the consumer, structured state should remain structured. Customer, transaction, policy and evidence should not have to become one enormous system message before a model can use them.

His description of system messages as “global variables” is provocative, but directionally right. Put every instruction everywhere and the behaviour becomes difficult to reason about. Pass only the relevant state into each bounded judgement and the system becomes easier to inspect, change and govern.

The unit economics only work in a narrow shape

“Decision AI is cheaper” is too broad to be useful.

It should be cheaper only when the workload has the right geometry:

Imagine one fraud case with a transaction, account history, device signals and merchant context. The business may need ten judgements: allow, challenge, hold, severity, likely pattern, reviewer queue, account linkage, evidence gap, next action and reason class.

A conventional design may repeatedly send the same evidence through a large generative model. A decision-shaped design may encode the evidence once and evaluate the bounded judgements together. That can reduce repeated context processing, decoding and malformed-output retries.

It also opens a new class of background use cases. Enterprises hold large amounts of “dark data”: documents, messages, tickets and operational records that are too expensive or too awkward to analyse continuously with large generative models. Cheap, bounded questions over shared state could make some of that data economically usable.

But “could” is doing a lot of work.

The savings disappear if evidence preparation dominates the cost, if every outcome needs a bespoke explanation, if latency is already acceptable, if the option taxonomy changes constantly, or if human review absorbs most cases. Vendor benchmarks are not a business case. The buyer needs a cost per completed, correctly handled case, including extraction, model calls, review, appeals, monitoring and failure.

Almeida uses “intelligence per dollar” as Jev’s north star. That is more useful than raw token price, but an enterprise still needs to define the numerator. Intelligence that does not improve an observed outcome is merely cheaper inference.

The right comparison is not price per token, but:

total cost per defensible outcome at the required service level

That number may favour a decision model. It may favour a smaller classifier, a rules engine, a generative model, or no AI at all.

A probability is not governance

Decision systems have a better object for governance than fluent prose: a distribution over permitted outcomes.

That is useful, but not sufficient.

A calibrated model should be right about 80% of the time on cases to which it assigns 80% probability. If that relationship holds on representative data, an enterprise can set thresholds, create abstention bands, plan review capacity and monitor drift.

But four things can still go wrong.

The model may be calibrated on the wrong population. Good performance on a benchmark says little about mortgage exceptions, vulnerable-customer calls or tomorrow’s fraud pattern.

The correct answer may be missing. A bounded menu forces probability somewhere. A confident distribution over the wrong set of choices is still wrong.

The policy may be illegitimate. Accurate predictions do not make an action lawful, fair or contestable.

The system may optimise the proxy. If the measured outcome is “case closed,” the model may learn to close cases, not solve customers’ problems.

So governance cannot stop at calibration. It needs three separate layers:

Keep those layers separate. Otherwise a model update silently becomes a policy change.

Almeida argues that safety for programmable AI should not rely on the model spontaneously refusing a request. The enterprise implication is not “remove safety.” It is the opposite: do not outsource safety to a hidden behaviour inside the model. Put it in an inspectable policy layer with explicit constraints and tests.

The menu is part of the model

Hume’s most operationally important finding is not speed. It is option sensitivity.

Adding an irrelevant option changed the odds between existing choices. Reversing option order moved one reported technical-support probability from roughly 0.84-0.89 to 0.93-0.96.

That means the choice menu is not harmless interface copy. It is model input.

A team can therefore change a production decision without touching the model, simply by reordering options, renaming a category or adding “other.” If 0.90 triggers an adverse action, a presentation change can become a policy event.

This should make every enterprise buyer slightly uncomfortable.

Treat option wording, order, taxonomy, thresholds and model version as one controlled decision surface. Test permutations. Add distractors. Remove options. Paraphrase labels. Repeat identical calls. Test missing evidence and out-of-domain cases. If the result crosses an action threshold under changes that should be irrelevant, the system is not stable enough for that action.

This is not a Jev-specific criticism. It is a first-principles requirement for any bounded-choice model.

Where Decision AI should start

The seductive pitch is to automate a high-value decision. The sensible move is the opposite.

Start where four conditions meet:

Ticket routing is better than credit approval. Document completeness is better than claim denial. Refund review support is better than automatic account closure.

Why? Because the first deployment is not really buying automation. It is buying evidence about whether automation deserves to exist.

A serious pilot should answer seven questions:

  1. What exactly is the decision? If the allowed outcomes cannot be stated cleanly, the process is not ready.
  2. What evidence may be used? Separate facts available at decision time from facts that leak in later.
  3. What is the cost of each error? False positives and false negatives rarely cost the same.
  4. What outcome will prove value? Human agreement is not enough. Measure what happened next.
  5. Where must the system abstain? Unknown, other and out-of-domain are production states, not edge cases.
  6. How stable is the decision surface? Attack wording, order, missing data and distribution shift.
  7. Who owns the action? Name the person or function that can change policy, stop the system and hear an appeal.

Run in shadow mode first. Then automate a narrow, reversible band. Widen it only when observed outcomes justify it.

For background agents, add one more rule: design the escalation path before the happy path. A system that works only when the model is confident has not solved the enterprise problem. It has hidden the difficult cases.

That is slower than a demo and faster than a scandal.

The architecture that follows

Once we stop asking one model to be the whole company, the enterprise stack becomes clearer.

These are different jobs, not competing eras.

A copilot can help an analyst understand a case. A decision model can estimate the path. A policy engine can decide whether automation is allowed. An agent can gather missing documents. A person can hear the appeal.

For some workflows, the right design may also be a cascade: a cheap bounded model handles clear cases, a larger model inspects ambiguous ones, and a person receives the consequential remainder. Calibration does not eliminate the human queue. It helps decide what belongs in it.

The mistake is not using generative AI. The mistake is using a product optimised to satisfy a person as if it were a system optimised to survive reality.

The deeper change is that the agent stops being the intelligence. It becomes the coordinator of several kinds of intelligence and control.

That is what turns an assistant into automation.

The question buyers should ask

Do not ask vendors only which foundation model they use.

Ask:

The last question is the test of whether this is engineering or theatre.

The next phase of enterprise AI will not be won by the model that sounds most certain. It will be won by systems that know when a decision is bounded, when uncertainty is actionable, when a person must stay in the loop and when the objective itself is wrong.

The real line is not between copilots and agents.

It is between intelligence designed to win approval and intelligence designed to run inside accountable software.

The first can help us work.

The second may finally let the work happen in the background.

Challenge

Pick one automated or near-automated decision in your business this week.

  1. Is the consumer a person or code?
  2. Can you decompose it into decisions small enough to test, with an explicit policy layer combining them?
  3. If the background dependency fails or refuses, what happens next?

If the answers live only in a demo script, you still have an assistant wearing an automation badge.

Sources

  1. Diogo Almeida, “Why I couldn’t build Jev at OpenAI”, Latent Space, 21 September 2026. The discussion supplies the distinction between human-facing assistants and models built for code; the “System One” and “large programmable model” framing; decomposition into small evaluable decisions; structured state; intelligence per dollar; and background software as the destination. These are Almeida’s arguments and product claims, not independent evidence that Jev delivers them for a buyer’s workflow.
  2. Diogo Almeida, “Jev CEO: I made ChatGPT, now I’m building what’s next”, AI Engineer, 31 July 2026. The talk supplies the distinction between assistance and automation, the criticism of preference optimisation for autonomous work, and the emphasis on verifiable rewards. Those claims are TypeSafe’s argument, not independent proof that any product solves the problem.
  3. Archer Hume, “Jev’s Architecture Unmasked”, 17 September 2026. Hume reports roughly 10,000 calls against Jev 1.13.0 from one early-access account and region. Shared computation, option interaction, batching and calibration results are observed behaviours or tests; the internal architecture remains inference.
  4. Evidence boundary: no public benchmark cited here proves Jev is calibrated, cheaper or safer for a buyer’s own workflow. Those claims require representative held-out outcomes, end-to-end cost measurement and production monitoring.

About the author

Prashant Chamarty helps enterprises put AI into real operating models: roles, workflows, decision rights, and controls, not just tools. Based in London. Writing and speaking on that gap.

Cite this page

Prashant Chamarty. (2026). The Most Dangerous AI Is the One That Does Exactly What We Asked. Retrieved from https://prashantchamarty.io/essays/the-most-dangerous-ai-is-the-one-that-does-exactly-what-we-asked

Interested in speaking engagements? Get in touch via email or LinkedIn.

Related essays

← Back to all essays