The Frontier Is Jagged. Your Workflow Needn't Be.

26 September 2026 · Prashant Chamarty · 5 min read

operating-model agents controls governance

TLDR

Models are uneven: they can solve a hard task and miss a simpler one. A harness cannot erase that limit. It can give an agent current context, restrict its tools, test the answer and route uncertain work to someone who owns the decision. Build those boundaries around a real workflow, then measure both the work completed and the exceptions created. The next problem is organisational: if nobody owns the boundaries, the harness becomes another layer of software nobody can safely change.

An agent can solve a difficult problem and miss an easy one. The mistake is to treat either result as a stable measure of what it can do next.

Dell’Acqua, Mollick and their co-authors called this the jagged technological frontier. In their study, AI helped consultants on tasks within its frontier and made them less likely to solve a selected task outside it. The boundary was not obvious from how difficult a task looked to a person.

A better model may move that boundary. It cannot map every boundary inside your business. The business has to specify the work, the evidence a decision needs and the cost of getting it wrong.

A recent survey describes the harness as the infrastructure around an agent: workflow, memory, skills, tools and recovery. BCG describes an operating system of context, rules, control and quality gates. I find the shift useful, provided we ask who may set and change those rules. The harness makes failure more visible and limits how far it can travel. It does not make the model infallible.

Start with a decision

Take an illustrative payment enquiry, not a client incident. A customer asks whether a transfer has settled. The agent reads a transaction record, drafts an answer and offers to update the case. In a clean test it finds the right record.

Now the ledger says pending, a risk system shows a hold, and the case note is from last week. Each source can produce a plausible answer. Which one establishes settlement? Who is allowed to tell the customer? What happens if the records disagree?

I would put a deterministic step before the model: retrieve the latest permitted records, keep the source and timestamp, check the customer identity, and stop the route when required records conflict. Let the model explain verified state and draft the next question. It should not turn a missing fact into a promise.

That design gives the model better ground. It also gives the organisation a place to say no.

Someone must own each boundary

For that one workflow, I would want five answers in writing:

  1. Context. Which records may be used for this purpose? How fresh must they be? Which source wins when they conflict?
  2. Tools. May the agent read, draft, open a case or change an account? Each action needs its own scope.
  3. Checks. What can a rule verify? When can a second model review help? Which decisions require a person? A model reviewing a model remains a technical check, not an accountable approver.
  4. Handoff. Who receives a stopped case, what evidence do they see, and when must they act? A queue without an owner is delayed work.
  5. Change. Who may widen access or relax a gate? What evidence would justify that change, and who can shut it off?

These are operating decisions. The engineering has to enforce them, but engineering cannot decide on its own which customer risk the business accepts. My earlier essay, “Done, Keys, and a Second Check,” focused on write rights and stop authority. Here the question starts earlier, with what the agent sees, and continues after its answer, with the people who inherit the edge cases.

What changes after the first deployment

The first-order result is straightforward: fresh retrieval and hard checks reduce avoidable mistakes in the selected workflow. The customer gets fewer answers based on stale or mismatched records. The agent can handle the cases that fit the rule.

The second-order effect is a change in human work. Routine cases leave the team; disputed and unusual cases stay. If the business measures only agent completion, it misses the time spent investigating exceptions. A larger queue may make the service slower even as the agent gets faster. Track the age of that queue, correction time, customer outcomes and the amount of work pushed to reviewers.

The third-order effect sits in the organisation’s ability to learn. When common cases are automated, newer staff see fewer of the examples from which they once learned. Meanwhile each new data source and tool increases the number of ways the workflow can fail. If the review team has no route to change the rules, the harness freezes yesterday’s assumptions into tomorrow’s operations. Keep a path for people handling exceptions to report patterns, revise the checks and test those revisions before expanding access.

Those effects are hypotheses to test in each deployment, not measured results from the illustrative case. A harness should earn its complexity. Start with one workflow. Test stale data, conflicting systems, wrong-customer retrieval and a tool result that tries to redirect the agent. Measure the whole job, including supervision and repair, before carrying the pattern to another team.

The frontier will remain jagged. A well-designed workflow does not have to wander across it blind.

Sources

About the author

Prashant Chamarty helps enterprises put AI into real operating models: roles, workflows, decision rights, and controls, not just tools. Based in London. Writing and speaking on that gap.

Cite this page

Prashant Chamarty. (2026). The Frontier Is Jagged. Your Workflow Needn't Be.. Retrieved from https://prashantchamarty.io/essays/the-frontier-is-jagged-your-workflow-neednt-be

Interested in speaking engagements? Get in touch via email or LinkedIn.

Related essays

← Back to all essays