A checklist for verifying AI output before it reaches a client, a filing or a decision. It is the differentiating control in this program and the practical companion to Hallucination and Verification.

The check exists because AI errors read fluently. A confident, well-formed answer does not announce itself as wrong, so verification has to be a step rather than a judgement about how the output feels.

Fluency is not evidence. The check that matters is whether someone opened the source.


The four checks

Run these four on any AI-assisted output before it is used.

1. Figures against source. Every number traced to the system of record or an identified external source — never to the model’s output, and never to the previous cycle’s document.

2. Claims against evidence. Every statement of fact checked, with particular attention to anything about a client, a person, a result or a comparison.

3. Citations opened. Every reference confirmed to exist, to say what it is claimed to say, and to be current. A citation that cannot be opened is removed, not reworded.

4. Commitments identified. Every promise, price, date, scope or service level confirmed as something the business can and intends to deliver.

Four checks, minutes to run on a familiar output type. The fourth is the one unique to AI-drafted communication, and it catches a class of error that accuracy checks miss.


By output type

Output Additional checks
Client report Consistency with the prior cycle; every figure reconciled to the accounts
Board or management pack Decisions page present; every exhibit ties to a source
Proposal Terms set by a human; capability claims confirmed with delivery; references permitted
Support response Answer against the knowledge base; escalation rules applied
Research briefing Source set bounded; limits of the sources stated
Marketing content Brand constraints applied; every factual claim substantiated
Financial commentary Figures from the frozen extract; variance causes confirmed

How to record the check

A check without a record is an intention. Four options, in ascending order of formality:

  • A tick in the checklist, retained with the output.
  • A version note naming the verifier and the date.
  • An approval in the system, where the workflow runs through one.
  • A line in the document itself, where the audience needs to see it.

The test: if the output were questioned six months later, could you show that it was checked, by whom, and when? Where the answer is no, the control does not exist in evidence.


What to do when a check fails

Stop and assess before the output proceeds.

  • Trace the failure to its source — a figure, a claim, a citation or a commitment.
  • Decide: correct the data, correct the output, or disclose the limitation.
  • Never override silently. An overridden check is a future correction.
  • Log the failure and its cause, because the log is what improves the workflow.

A failing check is the system working, not the system failing. That framing matters: where finding a problem carries a penalty, problems stop being reported, and unreported problems are the expensive kind.


Making it a habit

  • Put the check in the workflow, as a named step with an owner.
  • Practise it in training, on outputs containing a seeded fabricated citation and a plausible wrong figure.
  • Praise the catch. The person who finds a serious error should be thanked publicly.
  • Report the verification rate as one of the adoption measures.

Frequently asked questions

What should I check in AI output?

Figures against source, claims against evidence, every citation opened, and any commitment confirmed as deliverable.

How long does verification take?

Minutes on a familiar output type. Where it takes as long as the work, the fix is to narrow the AI task until the check is short — not to skip the check.

Why is verification necessary if the model is good?

Because reported hallucination rates remain material — roughly 22% to 94% depending on task and measurement, with one benchmark finding 13.6% of responses grounded (Stanford HAI, AI Index 2026) — and because fluent errors do not announce themselves.

Do we need to verify everything?

Everything that reaches a client, a filing or a decision. Internal drafts can carry a lighter check, provided the check is recorded where the output is used.

What if verification keeps finding the same error?

That is the workflow telling you something. Log it, fix the cause — the prompt, the input, the source or the training — and re-check. Repeated corrections are a signal, not a cost of doing business.

Should the verifier be a different person?

Where the output is internal, one person can do both. Where it is client-facing, regulated or financially material, a second person on the figures is worth the minutes.


Next step

Run the four checks on your next cycle, record it, and report the verification rate monthly. See Training Staff to Verify AI Output and Designing a Human-in-the-Loop Workflow, or book a team training session and we will run the drill with your team.


Sources

  • Stanford HAI, AI Index 2026, with Vectara’s hallucination leaderboard and OpenAI model documentation (2025–2026): hallucination rates reported between roughly 22% and 94% depending on task and measurement, with one benchmark finding 13.6% of responses grounded.

Figures are cited from their sources and dated. Where a source is a vendor benchmark, the sample size is stated where published.