Most leadership teams cannot answer a basic question: how much did the business really spend on AI last quarter?
Finance can identify the main software contracts. IT can list approved platforms. Function leaders can describe pilots. Employees may be paying for individual tools or using AI features bundled into wider subscriptions. Consultants, data work, integration, training and management time sit elsewhere again.
The result is not one AI budget. It is a trail of partially visible decisions.
That was manageable when AI activity consisted of a few experiments. It is not manageable when tools are embedded in sales, marketing, finance, customer service, operations and product development. Once spending is distributed, a business can increase its AI investment without making a deliberate capital-allocation decision.
The answer is not a bigger spreadsheet of licences. Leaders need an AI Investment Ledger: a single view of each material initiative, its full cost, the operating outcome it is meant to change, the evidence achieved and the next funding decision.
The ledger has five questions:
- What are we funding?
- Which business outcome should change?
- What is the full economic cost?
- What value has actually been realised?
- Should we stop, fix or scale it?
Used properly, it turns AI from a loosely connected set of activities into a governed portfolio of business investments.
The spending problem is broader than software
AI costs are easy to understate because the visible invoice is often the smallest part of the decision.
A customer-service tool may have a clear annual licence fee. Its real cost can also include data cleaning, integration with the customer platform, security review, prompt and workflow design, staff training, quality checks, parallel running and the manager time required to correct exceptions.
An internally built assistant may appear inexpensive because the model charges are low. The business may still be consuming scarce developer time, subject-matter expertise and leadership attention. A generative-AI feature within an existing suite may appear free even though adopting it creates process, control and training work.
The reverse problem also occurs. Leaders may treat all associated cost as an argument against the initiative, even when much of that spending creates a reusable capability. A clean customer-data layer, an approved model gateway or a reliable evaluation process may support several workflows rather than one.
The objective is not to make every cost look larger. It is to make the economics comparable.
For each initiative, separate eight cost categories:
- licences, model usage and infrastructure;
- external advisers, developers and implementation partners;
- data preparation, integration and systems changes;
- employee design, testing and subject-matter time;
- training, adoption and change support;
- risk, legal, security and quality assurance;
- parallel running, exception handling and rework; and
- ongoing ownership, maintenance and monitoring.
Record both cash cost and internal capacity. Cash determines affordability. Capacity shows what else the business is not doing.
This is particularly important in a smaller company. Ten senior employees each contributing a few hours a week can represent more economic cost than the software contract, even though none of it appears in the AI line of the management accounts.
The value problem starts with the wrong unit of analysis
Many businesses measure AI by tool: number of users, prompts, sessions, generated documents or hours reportedly saved.
Those measures can help diagnose adoption, but they do not establish business value. A tool can be heavily used while making little difference to revenue, margin, cash, risk or customer experience. It can also save time in one task while adding review, correction or coordination elsewhere.
The useful unit of analysis is the workflow.
Take proposal production. The workflow starts when an opportunity is qualified and ends when a commercially approved proposal reaches the customer. Drafting is only one step. If AI cuts writing time by 50 per cent but proposals still wait for pricing approval, legal review or senior sign-off, overall cycle time may barely move.
The business case should therefore describe an end-to-end change:
Reduce the median time from qualified opportunity to approved proposal from seven working days to four, while maintaining gross-margin discipline and reducing senior review time.
That statement can be tested. It identifies a workflow, a baseline, a target, a quality condition and a form of capacity value.
Compare it with: "Use AI to write proposals faster." The latter encourages activity. It does not define an investable result.
For every initiative, name one primary economic outcome. It should fall into one of four categories:
- Revenue: higher conversion, retention, price, volume or speed to revenue.
- Cost and capacity: lower external spend, reduced handling cost, fewer roles required for growth or capacity released for higher-value work.
- Quality and risk: fewer errors, stronger compliance, lower loss exposure or more consistent decisions.
- Strategic option: a capability that opens a new proposition, channel, data asset or business model.
Secondary benefits can be recorded, but they should not be used to rescue an initiative whose primary case is failing.
Time saved is not value until the business uses it
This is the most common error in AI business cases.
An initiative saves 20 minutes per task. The task occurs 3,000 times a year. The hourly employment cost is £30. The spreadsheet reports a £30,000 annual benefit.
But the payroll has not changed. Output has not increased. Customer response time is unchanged. The people affected are simply less busy for part of the day, or the saved time is absorbed by email, meetings and additional checking.
The saving is real as capacity, but it is not yet realised as financial value.
Every time-saving claim needs a conversion decision. Leadership should choose one of four routes:
- Remove cost. Eliminate overtime, temporary labour, contractor spend or a planned hire.
- Increase throughput. Process more orders, cases, campaigns or customer interactions with the same team.
- Improve service. Reduce response time, backlog, error or failure demand in a way customers value.
- Redeploy capacity. Move named hours and people to higher-value work with a clear output measure.
If none of these routes is chosen, classify the benefit as potential capacity, not realised value.
This distinction is not pedantic. It protects the leadership team from adding together dozens of theoretical savings that cannot all be captured. It also forces managers to design the operating change around the tool.
Build the ledger in five passes
The first version does not require perfect data. It requires consistent decisions.
Pass 1: Find the investment
Start with every material AI-related activity, whether approved centrally or initiated locally.
Search supplier records, expense claims, corporate cards, departmental budgets, development backlogs and transformation plans. Ask function leaders which tools, embedded features, pilots, automations and external partners they are using. Include initiatives that have been paused but still carry cost or dependency.
Avoid turning the exercise into a hunt for unauthorised users. If employees believe disclosure will lead to blame, the inventory will be incomplete. The purpose is to establish facts, remove unsafe duplication and make better funding decisions.
For each initiative, record:
- a plain-English name;
- the workflow and function affected;
- the executive sponsor and operational owner;
- current stage: explore, prove, deploy, scale or retire;
- the users or teams involved;
- key suppliers and dependencies; and
- the date of the next decision.
The inventory often reveals several versions of the same idea. Three functions may be testing meeting-summary tools. Two suppliers may be building similar knowledge assistants. Different teams may be paying for overlapping research or content products.
Duplication is not always waste. Parallel experiments can be useful while uncertainty is high. It becomes waste when nobody knows they are parallel or when temporary trials turn into permanent subscriptions without comparison.
Pass 2: Name the intended outcome
Every initiative needs an outcome that a business leader, not only a technical lead, is accountable for.
Use a one-sentence investment thesis:
We are investing £X over Y months to change Z workflow from its current baseline to the target outcome, while maintaining the stated quality and risk limits.
For example:
We are investing £60,000 over six months to reduce manual invoice exception handling from 1,200 to 500 cases a month, without increasing payment errors or supplier disputes.
The sentence exposes weak cases quickly. If the team cannot name the workflow, baseline or target, it is not ready for scaled funding.
Do not demand false precision from an early experiment. Exploration has value when it resolves a defined uncertainty. In that case, the outcome might be evidence: prove whether product descriptions can be generated within the company's accuracy and brand standards at a unit cost below an agreed threshold.
The key is to fund learning deliberately rather than describe an unbounded pilot as innovation.
Pass 3: Calculate the full cost
Build a 12-month view using the eight cost categories. Separate one-off implementation cost from recurring run cost. Where usage drives cost, show the unit economics and the assumption behind scale.
Also identify avoided cost. If the initiative replaces a planned software purchase, agency retainer or additional hire, make that explicit. Avoided cost is often more credible than a broad productivity estimate because it links to a decision the business would otherwise make.
Allocate shared foundations sensibly. Do not charge the entire data platform to the first use case, but do not pretend the platform is free. A simple method is to show shared enabling cost separately, then assign a reasonable share to each scaled workflow for portfolio comparison.
The result should answer two questions:
- What cash will leave the business if we continue?
- What scarce internal capacity will the initiative consume?
Pass 4: Test realised value
Value evidence should strengthen as funding increases.
At exploration stage, credible evidence might be technical feasibility, user demand or a validated baseline. At proof stage, it should include measured workflow performance with a representative sample. At deployment stage, it should show adoption, quality and actual operating change. At scale, it should connect to management accounts, customer outcomes or another auditable enterprise measure.
Use an evidence ladder:
- Claim: a benefit is expected.
- Observation: users report improvement.
- Measurement: a defined task or workflow metric changes.
- Attribution: the change can reasonably be linked to the initiative.
- Conversion: management turns the improvement into revenue, cost, capacity, quality or risk value.
- Sustainment: the result persists after the initial push.
An initiative should not be presented as proven when it has only reached observation. Equally, a valuable early test should not be rejected because it has not yet produced a full-year P&L result. The evidence standard should match the funding stage.
Pass 5: Make a funding decision
Every review should end with one of four decisions:
- Stop: the outcome is no longer important, the economics are unattractive, evidence is weak after a fair test or the risk cannot be justified.
- Fix: the opportunity remains sound, but a specific constraint in data, workflow, adoption, ownership or control must be resolved before more funding.
- Scale: evidence is strong enough to extend the workflow, user group or business area, with updated cost and control assumptions.
- Hold: the initiative remains valid, but timing or dependency means no further commitment should be made now.
"Continue" is not a decision. It usually means the leadership team has avoided making one.
Stopping is not an admission that the original choice was foolish. A disciplined portfolio expects some initiatives to fail or lose priority. The waste lies in allowing a weak initiative to survive because the sponsor is senior, the launch was visible or the spend is dispersed.
Use five rules to prevent manufactured ROI
AI returns are particularly easy to overstate because activity is measurable long before economic value is visible. Five rules help keep the ledger credible.
1. Freeze the baseline before the intervention
Record the current volume, time, cost, quality and service performance before the pilot changes behaviour. If no baseline exists, use a short measurement period before claiming improvement.
2. Count net change, not gross benefit
Subtract new review work, exception handling, vendor cost, model usage, quality failures and management overhead. A 30 per cent reduction in drafting time is not a 30 per cent workflow saving if review effort doubles.
3. Separate capacity from cash
Report hours released, cash removed and revenue created as different values. Do not convert every hour into payroll savings.
4. Name the conversion owner
If capacity is meant to increase sales activity, the commercial leader owns the conversion. If it is meant to avoid recruitment, the relevant executive and finance own that decision. Technology cannot capture the value alone.
5. Do not add benefits that compete for the same capacity
The same saved hour cannot reduce headcount, increase output and improve service simultaneously. Choose the primary conversion route and treat other possibilities as options, not booked value.
These rules make the business case more conservative. They also make it more useful. A credible modest return deserves more confidence than a large number assembled from incompatible assumptions.
Run a 30-day AI Investment Review
A leadership team can establish the first reliable ledger in a month.
Week 1: Discover
Create the portfolio inventory. Reconcile procurement, expenses, departmental budgets, technology plans and employee disclosures. Group duplicate or connected initiatives around the workflow they affect.
Output: one list with owners, stage, users, suppliers and next decision date.
Week 2: Reconstruct the economics
For the largest or most important initiatives, write the investment thesis, baseline, target and full-cost view. Identify shared foundations and the management action needed to convert capacity into value.
Output: a comparable one-page case for each priority initiative.
Week 3: Challenge the evidence
Interview users and process owners. Inspect samples, workflow measures, quality results, customer effects and finance data. Distinguish adoption from performance and potential value from realised value.
Output: an evidence rating and a short list of gaps that can genuinely change the decision.
Week 4: Reallocate
Hold a portfolio review led jointly by the CEO or business sponsor and finance, with technology, operations and risk input. Decide stop, fix, scale or hold. Set funding limits, milestones and the next review date.
Output: a funded portfolio, not merely an updated register.
The review should also identify two types of enterprise action. First, common foundations worth funding once, such as secure access, data integration, supplier standards or evaluation methods. Second, recurring duplication or unmanaged use that should be consolidated.
The governance should be light, but it must be real
SMEs do not need an elaborate AI committee for every tool. They do need clear ownership of investment decisions.
A practical model has four roles:
- Business sponsor: owns the outcome and decides whether the workflow matters.
- Finance partner: validates cost, baseline, value conversion and funding stage.
- Technology or data owner: confirms feasibility, architecture, security and supplier implications.
- Operational owner: changes the workflow, manages adoption and sustains performance.
Risk, legal, people or customer expertise should join where the use case requires it. Not every initiative needs every function in every meeting.
Set thresholds. A low-cost, reversible assistant used on non-sensitive work may follow an approved fast path. A customer-facing agent, employment decision tool or initiative with material data, regulatory or reputational exposure needs stronger evidence and controls.
The portfolio review cadence can be monthly while investment is moving quickly, then quarterly once the system is stable. The agenda should remain short:
- What has changed in cost, evidence or risk?
- Which milestone has been met or missed?
- What management action converted value?
- What should be stopped, fixed, scaled or held?
- Where should the next pound go?
Questions the board or leadership team should ask now
Use these ten questions as a first diagnostic:
- Can we state our total AI cash spend and internal capacity cost for the last quarter?
- How much of that spend sits outside technology or central transformation budgets?
- Which five workflows account for most of the investment?
- Does every material initiative have a named business sponsor and operational owner?
- Is the primary outcome stated as a measurable workflow or economic change?
- Which reported benefits are realised, and which are only potential capacity?
- What quality, risk or customer measures could invalidate the value claim?
- Which initiatives have passed their original decision date without a fresh review?
- What have we stopped in the past quarter, and where was the funding moved?
- Which common capability would make several valuable workflows easier to scale?
If the first question cannot be answered, do not begin with a debate about the next AI platform. Begin with the ledger.
From AI enthusiasm to investment discipline
The next phase of AI adoption will reward businesses that can do two things at once: move quickly enough to learn and apply enough discipline to concentrate resources.
That does not mean every experiment needs a detailed five-year return model. It means the purpose of the experiment must be clear, the full cost must be visible and the next decision must have a date.
An AI Investment Ledger creates that discipline. It gives founders, CEOs and CFOs a shared language for cost, evidence and value. It helps technology leaders distinguish foundations from local tools. It gives operational leaders responsibility for changing the workflow, not merely adopting software.
Most importantly, it makes stopping normal. Capital released from a weak initiative can fund a stronger workflow, better data, safer controls or the people needed to turn a promising test into an operating result.
Allington Advisors helps leadership teams assess AI portfolios, redesign priority workflows and build practical investment cases, governance and implementation plans. If your AI activity has grown faster than your visibility of cost and value, a focused investment review can establish the facts and identify where to stop, fix or scale.
