Measurement is what converts an AI program from a set of anecdotes into a set of decisions. It is also where most programs go wrong, in one of two directions: they measure nothing and cannot defend the spend, or they measure everything and nobody reads the numbers.

This guide sets out the four measures that matter, how to capture them, the measures that mislead, and how to report honestly. It is part of the Skills, Change & Measurement pillar.

The question is not whether AI is being used. It is whether the work is faster, no less accurate, and controlled.


The four measures

Four measures are enough at small-business scale, and each answers a different question.

Measure The question it answers How to capture it
Usage Is the workflow being run? Cycles completed by the team it was designed for
Time Is the work faster? Cycle time against the pre-change baseline, including verification
Quality Is the output accurate and consistent? Corrections per cycle, and the stage at which they are found
Risk Are the controls holding? Verification steps recorded; incidents and their causes

The pair that carries the most information in the first quarter is time and quality together. Time alone can improve by skipping the check; quality alone can improve by slowing everything down. Read together, they show whether the workflow is genuinely better.


Set the baseline first

A measurement introduced after a change is not a baseline, and it will not be believed.

  • Capture hours per cycle, end to end, including review.
  • Capture cycles per year, so the annual figure is real.
  • Capture the correction rate, and where corrections are currently found.
  • Capture who does the work, because the value depends on whose time is returned.

The baseline belongs to the person who does the work, not the person who wants the project. Where the two disagree, the practitioner’s number is the one to use.


Measures that mislead

Four measures look useful and frequently mislead.

  • Tool licenses issued. A license is a purchase, not an adoption. Usage of the workflow is the measure.
  • Prompts run. Volume of interaction says nothing about whether the output was used or correct.
  • Deflection rates. In support or content, “handled without a person” is only good if the issue was actually resolved — pair it with satisfaction.
  • Self-reported time saved. People are generous about time saved and imprecise about where it went. Measure cycle time instead.

Where any of these is the only measure on the dashboard, the program is being assessed on activity rather than outcome.


Reporting honestly

A one-page monthly report, with four lines, is enough. Three practices make it credible.

  • Report the baseline, so the comparison is visible rather than asserted.
  • Report the use cases that did not work, with the cause. A program that reports only wins stops being read.
  • Report the verification rate, because it is the measure that shows the controls are holding rather than assumed.

The test to apply before circulating: could a sceptical finance lead check these numbers themselves? Where the answer is no, the report is advocacy rather than measurement.


What to do with the numbers

Three decisions follow from measurement, and each should be taken deliberately.

  • Continue. The workflow is used, faster and no less accurate. Extend it to the next audience or cycle.
  • Adjust. One measure has not moved. Understand why before extending — usually the check is too slow, or the workflow adds steps that were not designed away.
  • Stop. The measure did not move, the cause is understood, and the use case is not viable at this scope. Record the lesson and move to the next candidate.

The third decision is the one that most programs avoid, and avoiding it is how a portfolio of half-working workflows accumulates. See The AI Use-Case Prioritization Matrix.


A worked monthly report

One page, four lines, and a baseline row.

Measure Baseline This month Trend
Usage 12 report cycles a year 2 of 2 cycles run through the workflow On plan
Cycle time 14 hours 10.5 hours including verification Improving
Corrections 3 per report, found at final read 1 per report, found at the review gate Improving
Verification recorded None 2 of 2 documents Control in place

Two lines of commentary sit beneath it: what the corrections were caused by, and what the next cycle will change. One line records the use case that was stopped, with its cause.

The report is deliberately dull, which is what makes it credible. A finance lead can check every number against the finance system, and nothing in it is an assertion about enthusiasm or culture.

What good looks like at 90 days

A program that is working has a recognizable signature.

  • The workflow runs without prompting, because it is part of the work rather than an extra step.
  • Cycle time has fallen, and the decline is attributable to drafting rather than to a skipped check.
  • Corrections are found at the gate, not by a client, and their number is falling.
  • Verification is recorded consistently, so the control exists as evidence.
  • A decision has been taken — continue, adjust or stop — on the basis of the numbers.

Where all five are true, the second use case is straightforward, because the baseline method, the measures and the review discipline are already in place.

Common mistakes

  • No baseline. The result cannot be shown, so the program is judged on impressions.
  • Usage as the only measure. Activity is not outcome.
  • Time without quality. Speed measured without accuracy is a warning, not a win.
  • Self-reported savings. Cycle time is measurable; memories are not.
  • Reporting only wins. The report stops being read, and the program loses its evidence base.
  • Measuring but not deciding. A dashboard nobody acts on changes nothing, and it quietly signals that the numbers are for reporting rather than for management.

Frequently asked questions

How do you measure AI adoption?

Measure usage of the workflow, cycle time against a pre-change baseline, corrections per cycle, and recorded verification steps. Four measures, reported monthly.

What is the best measure of AI ROI?

Payback against a measured baseline, with the cost lines and assumptions shown. The calculation is simple; the credibility comes from the baseline and the inclusion of verification time.

Why is usage a poor measure on its own?

Because a workflow can be used and still be slower, less accurate, or uncontrolled. Usage answers whether the work is happening, not whether it is better.

How do we measure quality for AI-assisted work?

Track corrections per cycle and where they are found. Corrections caught at the review gate indicate the workflow working; corrections found later indicate the gate failing.

How often should we report?

Monthly, on one page, and at every review point. More frequent reporting produces data nobody reads; less frequent reporting lets a problem run for a quarter.

What if we cannot measure the baseline?

Set one now, and be honest that the first period is a baseline rather than a result. The next cycle will produce a genuine comparison, which is better than none.

Who should own the measurement?

The program owner, with the use-case owner supplying the figures. Where the person measuring is also the person advocating, the numbers should be checkable by someone else.

Should the measures be shared with the team?

Yes, and before the pilot begins, so everyone knows what is being assessed and why. Unexplained measurement is read as surveillance, and it changes behavior in unhelpful ways.

What if a workflow improves one measure and worsens another?

That is the most common early pattern: time falls and corrections rise, because the check is finding errors the old process missed. Hold both measures for a second cycle before drawing a conclusion.

What is the right cadence for the review?

Monthly for the measures and ninety days for the continue, adjust or stop decision. A monthly decision is too frequent to be meaningful; an annual one lets a failing workflow run for a year.

Should the measures differ by workflow?

The four measures stay the same; only the baseline changes. Keeping the measures constant is what makes the portfolio comparable and the reporting sustainable.


Next step

Capture the four measures, set the baseline before the next change, and report monthly on one page. Run your own numbers with the AI Adoption ROI Calculator, see Training Your Team to Use AI Well, or book a team training session and we will set up the measurement with you.


Sources

  • Measurement framework reflects standard program-measurement practice applied to AI adoption: baseline before change, usage, time, quality and risk measures, and review-driven decisions.

No statistic in this article is invented; where figures appear in the linked guides, they are cited there with their source and date.