Back to all articlesoperations ai

AI Decision Support Systems: Which Decisions to Delegate, Which to Surface, and Which to Keep Human

A calibration framework for operations leaders deciding how much authority to give an AI agent. Three questions, three levels, and the three failure modes that follow from getting the level wrong.

AI Decision Support Systems: Which Decisions to Delegate, Which to Surface, and Which to Keep Human

You have stopped asking whether AI belongs in your operation. What you cannot answer is how much authority to give it.

The demos do not help. One vendor shows an agent that recommends a carrier and waits. Another shows an agent that books the carrier, notifies the warehouse, and updates the ledger before anyone opens a laptop. Both look competent. Neither tells you which of your decisions belongs in which bucket, and that is the only question that matters, because the two ways of getting it wrong are both expensive.

Give away authority over a decision that carries regulatory or financial weight, and you find out in an audit. Route everything through human approval instead, and you have built a queue that costs more attention than the manual process did. The agent saves nobody any time; it just moved the work from doing to reviewing.

This is what we mean by AI decision support systems: not a dashboard that suggests things, but a deliberate assignment of each decision to a level of autonomy. Getting that assignment right is the work. Most vendors skip it because only someone who has run the operation can do it credibly.


The Three Levels

Every operational decision an agent touches sits at one of three levels. The definitions have to be sharp, because the whole framework depends on the boundaries.

Delegate. The agent decides and acts. No human sees it before it happens. A human can see it afterwards, and the system records what was decided and why. Example: reordering a consumable that is below its reorder point from the contracted supplier at the contracted price.

Surface. The agent decides, prepares the action completely, and waits for a human to release it. The human's job is to approve or correct — not to redo the analysis. This distinction matters more than anything else in this article. If your reviewer has to reconstruct the decision to check it, you have not built surface; you have built a slower manual process with extra steps. Example: a vendor payment above a threshold, presented with the invoice, the PO match, and the payment history in one view.

Keep human. The agent may retrieve, summarize, and model, but the decision is made by a person and the agent never proposes a specific action as the default. Example: terminating a supplier relationship, or anything where the inputs that matter are not in your systems.

The mistake we see most often is treating this as a maturity ladder — start everything at keep-human, graduate to surface, eventually delegate. It is not a ladder. Some decisions should be delegated on day one and some should never move. The level is a property of the decision, not of your comfort.


The Calibration Test

Three questions, asked about one decision at a time. Score each 1 to 3.

1. How reversible is it?

If the decision turns out wrong, what does it cost to undo?

  • 3 — Cheap to reverse. Cancel the order, re-route the truck, adjust the forecast. Minutes, no counterparty involved.
  • 2 — Reversible with friction. A counterparty has to agree, or the reversal is visible to a customer.
  • 1 — Effectively irreversible. Money has left, data has been shared, a legal notice has been sent, someone has been told.

Reversibility is the single strongest predictor of the right level, because a reversible decision made badly is a small cost paid quickly, and that is exactly the kind of mistake a system should be allowed to make.

2. What is the blast radius?

If this decision is wrong in the same way 500 times before anyone notices, what happens?

  • 3 — Contained. One order, one shipment, one ticket. Errors do not compound.
  • 2 — Correlated. A pricing rule or routing preference that applies across a class of transactions.
  • 1 — Systemic. Touches compliance posture, safety, a regulated calculation, or every customer at once.

Blast radius is where the delegate level usually goes wrong. Teams evaluate a single instance of the decision, conclude it is low-stakes, and delegate it — without noticing that the agent will make it continuously and that a bad rule applies itself perfectly.

3. How often does a human currently override?

Pull the last 200 instances. How many times did a person change what the standing process or the existing rules produced?

  • 3 — Under 5%. The decision is already effectively mechanical, and your humans are rubber-stamping.
  • 2 — 5% to 20%. There is real judgment, but it is exercised on a minority of cases.
  • 1 — Over 20%. The formal process does not capture how the decision is actually made. Something important lives in people's heads.

This is the question teams skip, and it is the most informative of the three, because it is measured rather than imagined. A high override rate on a decision you were about to delegate is a warning that the rules you would encode are not the rules being followed.

Scoring

Add the three scores.

TotalLevelWhy
8-9DelegateReversible, contained, already mechanical
5-7SurfaceReal judgment or real consequence — a human release is worth its cost
3-4Keep humanIrreversible or systemic, and judgment is doing the work

One override: any decision scoring 1 on reversibility and 1 on blast radius stays human regardless of total. Irreversible and systemic together is not a place to find out your calibration was off.


A Worked Sort

Here is a real decision list from a mid-market distributor, scored on the three questions. It took their operations lead about ninety minutes.

DecisionRev.BlastOverrideTotalLevel
Reorder consumable below reorder point3339Delegate
Assign inbound ticket to a queue3339Delegate
Pick carrier for a standard domestic lane3238Delegate
Flag an invoice as a duplicate3227Surface
Release a payment under $10k1337Surface
Change a customer's credit limit2226Surface
Apply a discount outside the standard band2125Surface
Write off a receivable1214Keep human
Put a supplier on hold1113Keep human

Two things surfaced immediately. Carrier selection scored an 8 and was being manually approved on every load — pure queue cost, no judgment being added. And releasing payments under $10,000 scored a 7 rather than the 9 the finance lead expected, because irreversibility drags it down no matter how routine it feels. It stayed at surface, with the review reduced to a batch summary rather than line-by-line.

Do this with twenty of your own decisions before you evaluate a single vendor. The list is the specification.


The Three Failure Modes

Each comes from getting the level wrong in a specific direction.

Failure 1: Delegating an irreversible decision

The agent does exactly what it was told, continuously, and the damage is done before the first review cycle. This is the failure everyone fears and the least common in practice, because the decisions that are obviously irreversible get scrutiny.

The dangerous version is subtler: a decision that is reversible in the system but not in the relationship. Cancelling an order is a database update and a phone call to a supplier who now trusts you less. Score reversibility against the world, not against your ERP.

Failure 2: Surfacing everything

More common and more corrosive. Every decision routes to an approval queue, the queue grows faster than anyone can review it with attention, and approvals become reflexive. Now you have the cost of the queue and none of the control, because a reviewer who approves 300 items an hour is not reviewing.

The tell is approval latency dropping while volume rises. That is not efficiency. That is the control layer quietly turning into a rubber stamp — and unlike a missing control, this one produces an audit trail that says someone checked.

Failure 3: Keeping human what is already mechanical

The quiet one. A decision sits at keep-human because it always has, the override rate is 2%, and a person spends four hours a week producing the same answer the rules would. No incident ever results, so nothing forces a review. The cost is invisible and permanent, which is why the override-rate question is worth the effort of actually measuring rather than estimating.


Levels Move, But Only With Evidence

Calibration is not a one-time exercise, and the movement is not automatic.

A decision earns promotion from surface to delegate when the agent's proposals have been accepted without modification at a high rate over a meaningful sample — our rule of thumb is 95% acceptance over at least 200 instances, with the disagreements reviewed rather than counted. If your reviewers are correcting 1 in 10, the agent has not learned the decision; it has learned most of it, and the remainder is where your money is.

Demotion should be faster than promotion. One incident traceable to a delegated decision moves it back to surface immediately, and it re-earns delegation the same way it earned it the first time. Asymmetry here is deliberate: the cost of an unnecessary review is bounded, and the cost of an unwatched systematic error is not.

Review the whole map quarterly. Decisions drift — a supplier base consolidates, a regulation changes, a product line starts carrying different risk — and a level that was right in March can be wrong by September.


What to Do This Week

  1. List twenty decisions your team makes repeatedly. Not processes — decisions, each with an owner and an outcome.
  2. Pull the override data for the five you would most like to automate. Actual counts from the last 200 instances, not recollection.
  3. Score all twenty on reversibility, blast radius, and override rate.
  4. Start with your highest scorers, not your biggest pain. The 8s and 9s are where an agent earns trust cheaply, and trust is what funds the harder ones.
  5. Write down the demotion triggers before anything goes live. What specifically sends this decision back a level?

We do this calibration work with operators across factory and facility operations, back-office finance, customer success, and CS calling operations. The method is the same each time, and the map is different every time. If you want to see how it plays out on a floor rather than on paper, our write-up of warehouse autonomy calibration walks a full operation end to end. If you would rather work through your own list with us, book a session.


FAQ

What is an AI decision support system?

It is a system that assigns each operational decision to a level of autonomy — delegate, surface, or keep human — and enforces that assignment. The distinguishing feature is not that AI produces a recommendation, but that the authority boundary for each decision class is deliberately set, recorded, and reviewed rather than left implicit.

How do I decide which decisions an AI agent can make on its own?

Score each decision on three questions: how reversible it is, how large the blast radius is if it is wrong repeatedly, and how often a human currently overrides the standing process. Decisions that are cheap to reverse, contained in effect, and already overridden less than 5% of the time can be delegated. Anything irreversible and systemic stays human.

What is the difference between human-in-the-loop and full automation?

Human-in-the-loop — what we call surface — means the agent completes the analysis and prepares the action, but a person releases it. Full automation, or delegate, means the agent acts and the human reviews after the fact if at all. The practical test is whether a reviewer can approve without reconstructing the decision. If they cannot, the loop is costing more than it protects.

How many decisions should be fully delegated?

There is no target ratio, and treating one as a goal produces bad calibration. In the operations we have mapped, the delegate bucket is usually the largest by transaction volume and the smallest by decision value — many routine, contained, high-frequency decisions, and very few consequential ones.

When should an AI agent never decide?

When the decision is effectively irreversible and its blast radius is systemic at the same time — regulatory filings, safety interventions, terminating a relationship, anything where the inputs that matter are not represented in your systems. These stay human even when the agent could produce a defensible answer.

Need help with AI implementation?

We build production AI systems that actually ship. Not demos, not POCs—real systems that run your business.

Get in Touch