Verification is the skill that makes every AI-assisted workflow safe, and it is the one most often described rather than taught. A team told to “check the output carefully” has been given a value, not a process. A team that has practised catching a fabricated citation in a session has a habit.

This guide covers how to teach verification: what the check involves, how to drill it, how to record it, and why the drill matters more than the explanation. It is part of the Skills, Change & Measurement pillar.

Verification is a practised behavior, not a stated expectation. If it has not been drilled, it has not been taught.


What the check actually involves

Verification is four specific actions, not a general review.

  • Figures against source. Every number traced to the system of record or an identified external source, never to the model’s output and never to the previous cycle’s document.
  • Claims against evidence. Every statement of fact checked, with particular attention to anything about a client, a person or a result.
  • Citations opened. Every reference confirmed to exist and to say what it is claimed to say.
  • Commitments identified. Any promise, price, date or scope confirmed as something the business can and intends to deliver.

The four are short to state and require practice to apply consistently. The third is the one that catches the most damaging errors, because a fabricated citation attached to a real publication is very hard to spot without opening it.


Why it has to be drilled

Two reasons, and both are practical.

Fluency defeats intuition. AI errors read well. The usual signals that prompt a second look — awkward phrasing, a gap in the logic — are absent, so the check has to be a step rather than a feeling.

The error rates justify it. Reported hallucination rates in large language models range widely — roughly 22% to 94% depending on the task and measurement method, with one benchmark finding 13.6% of responses grounded (Stanford HAI, AI Index 2026, with Vectara’s hallucination leaderboard and OpenAI model documentation). The spread is wide because measurement is hard; the practical implication is narrow and firm: verification is not optional.


The drill

A verification drill has three exercises, run in one session.

1. The seeded citation. An output containing a fabricated reference, formatted to look real, attributed to a plausible publication. The learner must identify it. The lesson is that a citation looking right is not evidence that it exists.

2. The plausible wrong figure. A number in the correct range and format, but wrong, and inconsistent with a source the learner can check. The lesson is that figures must be traced, not judged.

3. The embedded commitment. A drafted message that promises a timeline or an outcome the business has not agreed. The lesson is that commitments in AI-drafted communication are a class of risk distinct from accuracy.

Each exercise takes minutes, and each produces a memorable example. A learner who has caught all three will check the fourth case automatically.


Recording the check

A check without a record is an intention. Four ways to record it, in ascending order of formality.

  • A tick in a checklist retained with the output.
  • A version note naming the verifier and the date.
  • An approval in the workflow system.
  • A line in the report or document itself, where the audience needs to see it.

The test: if the output were questioned six months later, could you show that it was checked, by whom and when? Where the answer is no, the control does not exist in evidence.


Making it stick

Three practices turn the drill into a habit.

  • Make the check visible in the workflow, as a named step with an owner. A step that is not in the process is not done.
  • Praise the catch. The first person to find a serious error in AI output should be thanked publicly, because the behavior the business wants repeated is the behavior it recognizes.
  • Report the verification rate as one of the adoption measures, so the check is managed rather than assumed.

The counter-productive alternative is a warning about hallucination with no practice, which produces people who know the word and not the behavior.


A worked drill session

Sixty minutes with a team of six, run once, for one workflow.

  • Minutes 0-5: set-up. No lecture. The instruction is simply that the team will be given three outputs to check, and that all three contain a problem.
  • Minutes 5-15: the seeded citation. An output citing a named publication and year, with a figure attributed to it. The publication exists; the figure is from a different study. Most of the team find the citation, and the discussion is about how convincing it looked.
  • Minutes 15-30: the plausible wrong figure. A number within the expected range, formatted correctly, inconsistent with the source the team can access. The lesson: figures are traced, not judged.
  • Minutes 30-45: the embedded commitment. A drafted client email promising a delivery date the business has not agreed. This one catches people who found the first two, because it is an accuracy question in a different frame.
  • Minutes 45-60: the team’s own check. The team writes, or rewrites, the check step for their workflow, based on what they just did.

The session ends with the team owning the check rather than receiving it. That is the difference between a drill and a briefing.

Common mistakes

  • Describing verification rather than practising it. Knowledge does not become habit without a drill.
  • A check with no owner. “The team checks it” means nobody does.
  • No record. The control exists as practice and not as evidence.
  • Only checking figures. Claims, citations and commitments are separate classes of error.
  • Punishing the error rather than the missed check. People conceal mistakes they expect to be penalised for.
  • Never re-drilling. New staff, new tools and new output types all create new failure modes, so the drill has a shelf life.

Frequently asked questions

How do you teach staff to verify AI output?

With a drill, not a lecture: seeded citations, plausible wrong figures and embedded commitments, run in one session and repeated for new staff and new workflows.

What should staff check in AI output?

Figures against source, claims against evidence, every citation opened, and any commitment confirmed as something the business can deliver.

Why is verification so important?

Because AI errors read fluently, so ordinary judgement does not catch them, and reported hallucination rates are high enough that unverified output is unreliable as a rule.

How do we know verification is happening?

Record it — a tick, a version note or an approval — and report the verification rate as one of the adoption measures. Without a record, the control exists only as an intention.

Should we punish staff who miss an error?

Punish the missed check, not the error, and even then only where the check was skipped. A culture that penalises discovering mistakes produces concealment, which is far more expensive.

How often should the drill be repeated?

For new staff, for each new workflow, and annually as a refresher. New tools and output types create new failure modes, so the drill is never one-and-done.

Does the check slow the work down?

It adds minutes and prevents the kind of error that costs far more. Where the check takes as long as the work, the fix is to narrow the AI task until the check is short — not to skip the check.

What if the team finds no errors in the drill?

Then the seeded errors were too obvious. Make them plausible — a real-looking citation and a number in the right range — because the point of the drill is to show how convincing an error can look.


Next step

Run the three-exercise drill with the next team you train, add a record step to the workflow, and report the verification rate monthly. Use the Human-Verification Checklist, see Designing a Human-in-the-Loop Workflow, or book a team training session and we will deliver the drill with your team.


Sources

  • Stanford HAI, AI Index 2026, with Vectara’s hallucination leaderboard and OpenAI model documentation (2025–2026): hallucination rates reported between roughly 22% and 94% depending on task and measurement, with one benchmark finding 13.6% of responses grounded.

Figures are cited from their sources and dated. Where a source is a vendor benchmark, the sample size is stated where published.