The difference between suggesting and acting

A chatbot that writes a poor email wastes a minute of someone's time. An agent that sends it does not.

That single difference, between software that suggests and software that acts, is the line the mid-market is now being asked to cross. Most businesses have spent the last two years with assistive AI, where a person always sits between the model and the world. The tool drafts, summarises or proposes, and a human reads it, judges it and decides what to do. The worst case is a bad suggestion that someone catches.

An agent removes that person from the loop. It does not propose the refund, it issues it. It does not draft the reorder, it places it. It does not flag the price change, it makes it. The technology may be only marginally cleverer than the assistant you already use. The exposure is entirely different, because a wrong action now happens at machine speed, can repeat before anyone notices, and has no one standing between it and your customer, your bank balance or your data.

The right question about an agent is therefore not the one the demos answer. It is not "is it clever enough?" It is "how much authority should it have, and who is accountable when it is wrong?"

The gap between appetite and control

The direction of travel is not in doubt. Around two-thirds of companies are already exploring AI agents, according to BCG, and Deloitte expects up to three-quarters of companies to be investing in them by the end of 2026. Gartner expects 40% of enterprise applications to have task-specific agents embedded by the end of this year, against fewer than one in twenty a year ago.

What has not kept pace is control. Gartner predicts that more than 40% of agentic AI projects will be cancelled by the end of 2027, and is blunt that the reason will not be the technology. It will be cost, unclear value and weak risk controls. A Forrester survey found that 71% of firms deploying agents have no formal governance framework for them, while 64% intend to give their agents more autonomy within the year. McKinsey reports that 80% of organisations have already seen agents behave in risky ways, from exposing data they should not have touched to reaching into systems they should not have accessed.

Read those numbers together and the picture is clear. The appetite to let software act is running well ahead of the discipline to control what it does. That gap is where the money and the reputational damage are going to be lost.

There is a second trap laid specifically for the mid-market. Much of what is sold as "agentic" is not. Gartner has a name for it, "agent washing", the rebranding of ordinary chatbots and automation scripts as agents, and estimates that only around 130 of the thousands of agentic vendors in the market are the real thing. The first act of control is therefore a plain question to any supplier: what does this actually do without a human, and where is the evidence?

Authority is the decision, not capability

The instinct in most businesses is to judge an agent on what it can do. That is the wrong axis. What matters commercially is what it is allowed to do without you, because that is what determines both the value it can create and the damage it can cause.

It helps to stop thinking in terms of "using an agent" and start thinking in terms of levels of authority. There are four worth naming.

Level one, Assist. The agent drafts, summarises or suggests. A human does everything of consequence. This is the copilot most teams already have. Low value, low risk.

Level two, Recommend. The agent proposes a specific action and shows its reasoning, and a human approves each one before anything happens. Useful for building trust and gathering evidence before you go further.

Level three, Act with approval. The agent executes, but only through a defined gate. Routine actions may pass automatically within tight limits, while anything above a set threshold stops for a human decision.

Level four, Act autonomously. The agent executes within set boundaries without asking first. Humans supervise after the event, review by exception and step in when something looks wrong.

Most of the real prize sits at levels three and four, because that is where work is genuinely removed rather than merely sped up. It is also where the risk concentrates. The discipline is not to pick a level for "our AI" in general. It is to match the level to each individual task, and to make the agent earn its way up rather than starting at the top because a vendor said it could.

The Delegation Test

Before you grant a task to an agent, run it through six questions. Score each from 1 (low concern) to 5 (high concern). The rule that matters is at the end, and it is not an average.

  1. Reversibility. If the agent gets this wrong, can we undo it cheaply and quickly? Deleting records, moving money and emailing a customer are hard to unwind. Drafting an internal note is trivial to correct.
  2. Blast radius. How much damage can a single wrong action do, and how fast can it repeat at scale before a human notices? An error that touches one internal report is not the error that mis-prices a thousand orders overnight.
  3. Rule clarity. Are the rules explicit, stable and reasonably complete, or does the task really run on judgement, context and constant exceptions? Agents are strong where the logic is clear and weak where it is tacit.
  4. Data quality. Does the agent have reliable, current and complete information to act on, or is it working from partial, stale or messy data? An agent acting confidently on bad inputs is worse than no agent at all.
  5. Auditability. Can we see exactly what the agent did and why, and reconstruct it afterwards for a customer, an auditor or a regulator? If you cannot explain an action, you cannot defend it.
  6. Exposure. What is the regulatory, contractual and reputational cost if this goes wrong? Regulated advice, pricing, safety, anything touching personal data, all raise the bar sharply.

Now the rule. Do not average the scores. Your weakest dimension sets the ceiling. A task that is beautifully rule-based and data-rich but hard to reverse and high in blast radius is still a task you keep on a tight leash, because the day it is wrong, the clarity will not save you. One serious answer caps how far the task can be delegated, however comfortable the others feel.

Reading the result

The six questions collapse, usefully, onto two axes. On one, how costly it is to be wrong, combining reversibility and blast radius and exposure. On the other, how clear the task is, combining rule clarity and data quality. That gives four practical zones.

Low cost, high clarity: Automate. Repetitive, rules-based, low-stakes work is where agents earn their keep with least risk. Categorising invoices, routing inbound queries, reconciling routine transactions, chasing standard documents. Give these to an agent at level four and reclaim the time.

Low cost, low clarity: Monitor closely. Work that is low-stakes but genuinely fuzzy, such as first-draft responses or triage, can run autonomously, but with heavy logging and a human sampling the output regularly until the quality is proven.

High cost, high clarity: Approve each action. Where the rules are clear but a mistake is expensive, keep a human gate. Refunds above a threshold, price changes within set guardrails, payments over a limit. The agent prepares and executes the routine, and stops for sign-off on anything that matters.

High cost, low clarity: Keep human-led. Where stakes are high and judgement is central, the agent assists and recommends only. Hiring and credit decisions, anything regulated or safety-critical, sensitive communications with key customers. Here the human decides and the agent informs, not the other way round.

The most expensive mistakes come from treating a top-right or bottom-left task as if it were bottom-right, and letting an agent act autonomously on something high-stakes because a demonstration looked impressive.

The seven controls before you let an agent act

Deciding a task can be delegated to level three or four is not the same as being ready to do it. Before any agent is allowed to act, seven controls should be in place. None of them is exotic, and the absence of any one is a reason to wait.

  1. A named accountable human. One person owns the outcome of what the agent does, not merely its setup. "Which agent acted under whose authority" is becoming a formal question. Singapore's model governance framework for agentic AI, published in January 2026, already requires each agent to carry a verifiable identity and an audit trail of exactly that. Your version can be simpler, but the name cannot be blank.
  2. Hard limits on scope and spend. The agent can act only within explicit boundaries: which systems it may touch, which accounts, what maximum value, and what it must never do under any circumstances.
  3. An exception threshold. Anything above a defined size, novelty or uncertainty is escalated to a human rather than decided by the agent. Most damage happens at the edges, so route the edges to a person.
  4. A full audit trail. Every action logged, with the reasoning and the data behind it, so any decision can be reconstructed and explained after the fact.
  5. A kill switch. One deliberate, well-understood way to stop the agent immediately, tested before you need it rather than discovered in a crisis.
  6. A review cadence. Someone reads a sample of the agent's actions on a set rhythm and watches the error rate over time, not just once at launch. Agents drift as the world around them changes.
  7. An earn-your-way-up rule. Agents start low on the ladder and are promoted only when the logs show stable, accurate performance. This mirrors what the most careful adopters now do: begin in assisted mode and grant more autonomy only once the evidence supports it. Autonomy is earned, not granted on day one.

Where this connects to the profit

It is worth being clear about why this matters commercially, and not only as a matter of safety. In earlier insights we argued that most AI never reaches the profit and loss account because time saved is rarely converted into anything: the hours freed up are absorbed rather than banked. Agents are the first tool that can genuinely remove work rather than merely accelerate it, which is exactly what closes that gap. An agent that fully handles invoice matching does not make the task faster, it takes the task away.

That is precisely why the delegation decision is a commercial one, not an IT one. The value of an agent comes from operating at levels three and four, where work is actually removed. So does the risk. Getting the authority right is what lets you capture the first without being caught by the second. Businesses that delegate deliberately will convert agents into margin. Businesses that either refuse to let anything act, or let everything act at once, will spend 2026 and 2027 in the group Gartner expects to cancel their projects.

Where to start

You do not need an agent strategy. You need two or three well-chosen tasks.

Pick your most repetitive, rules-based, low-stakes processes, the clear "automate" candidates. Run each through the Delegation Test. Put the seven controls in place before anything goes live. Start the agent at "act with approval", watch the logs, and promote it to autonomous only once it has earned the trust. Treat any vendor claim of full autonomy as a prompt to ask what the tool really does without a human, and to see the evidence.

Above all, resist the two failure modes at either end. Handing a high-stakes, judgement-heavy process to an agent because the demonstration was slick is how the expensive mistakes happen. Refusing to let anything act at all is how you leave the one real productivity unlock of this cycle sitting on the table while your competitors take it. The advantage goes to the leaders who can tell the difference, task by task.

If you are being pitched agents and are not sure how far to trust them, that uncertainty is exactly what a short review resolves. Book an AI Agent Readiness Review with Allington Advisors and we will help you decide which tasks are safe to delegate, at what level, and what controls to put in place before you let anything act.