A hallucination is a confident, fluent, false statement. In a report, it is a liability: a fabricated figure, a misstated trend, or a cause that never happened. Readers act on reports, and a single fabricated number does not just undermine the page – it undermines trust in every figure the function produces.
This guide explains what hallucination risk looks like in reports, what the measured data says, where errors hide, and the verification discipline that catches them before release. It is the companion to Using AI to Draft Reports.
In a report, the danger is not that the AI is wrong. It is that the AI is wrong and confident, and nobody checks.
What hallucination is
A model does not “know” your revenue, your headcount or your margin. It predicts plausible text. Most of the time the prediction is right; on specifics – figures, trends, causes, names – it can be confidently wrong. This is not a bug a better prompt eliminates; it is a property of how the systems are built.
For reports, the consequence is simple: anything the model states as fact must be verified against a source before it is published.
What the measured data says
The problem has not gone away, and it is worst on exactly the content reports depend on.
| Benchmark | What it measures | Headline finding | Source |
|---|---|---|---|
| AA-Omniscience | Guessing vs. abstaining on hard factual questions | 22-94% hallucination across 26 top models | Artificial Analysis / Stanford AI Index, 2026 |
| Vectara Hallucination Leaderboard | Fabrication when summarizing a given document | Best model still fabricated in 13.6% of responses | Vectara, 2026 |
| OpenAI PersonQA | Hallucination on questions about people | o3: 33%; o4-mini: 48% | OpenAI o3 / o4-mini system card, 2025 |
| Stanford RegLab (legal) | Hallucination on legal queries | 69-88% on specific queries | Stanford RegLab |
The pattern: rates fall when the model is given a source to summarize, but they do not fall to zero. “Grounded” means “safer, still needing a check,” not “safe.”
Why reports are high-stakes
Reports concentrate the risk for three reasons:
- Every figure is a claim the reader may act on.
- Readers treat accuracy as a proxy for reliability, so one wrong number casts doubt on the rest.
- There is no quiet correction. A report that circulated with a wrong figure must be corrected publicly.
That is why verification is not optional. The cost of one unverified figure can exceed the entire saving from using AI in the first place.
Where errors hide
Hallucination is most dangerous when it hides in the specific, checkable details – the things a hurried reader assumes are right.
- Figures and totals – revenue, margin, headcount, percentages.
- Trends – “up 12% quarter on quarter” when it is not.
- Causes – an explanation the data does not support.
- Periods – a figure attributed to the wrong window.
- Entities and names – the wrong client, product or person.
- Comparisons – against the wrong baseline or target.
These are precisely the details that build credibility when correct and destroy it when wrong – and the details a model is most likely to fabricate, because they are specific and low-frequency.
How to catch it
The catch is a discipline: verify every figure and cause against a source of truth.
- Maintain a source of truth – the system of record and the metric dictionary. See Data Quality for Reporting.
- Verify against the source, not another model. Two AI outputs agreeing proves nothing.
- Flag, do not soften. Unverifiable claims are marked for resolution, not quietly reworded.
- Separate drafting from checking. The verifier should not be the drafter.
- Log it for high-stakes reports. Record the figure, its source and the verifier.
This is stage four of the human-verified workflow, and it is what makes AI safe in reporting.
A verification checklist
Run this before any AI-assisted content is published:
- Every figure traces to the system of record, with an as-of date.
- Every trend is confirmed against the actual prior period.
- Every cause is confirmed by someone who knows.
- Every period label matches the data window.
- Every entity or name is correct.
- Every comparison uses the right baseline.
- A named human has approved the report.
If any line fails, the content is not ready.
How hallucination shows up in different report types
The risk profile varies by report.
| Report type | Where errors land | Consequence |
|---|---|---|
| Board pack | Headline figures, trends | Leadership decisions on wrong numbers |
| Client report | Results, comparisons | Reputational damage |
| Investor update | Performance, outlook | Regulatory and trust exposure |
| Compliance report | Statements against rules | Penalty, non-compliance |
| Internal status | Progress, blockers | Wrong priorities |
Compliance and investor reporting carry the highest consequences, because the figures are relied on externally. That is where verification should be strictest, and where an AI-assisted draft should never reach the reader unchecked.
Building a verification habit
Verification works when it is a habit, not a heroic effort.
- Verify against the frozen snapshot, not the live system, so the check is against a fixed target.
- Check figures first, causes second, because a wrong figure invalidates everything after it.
- Keep a running source log as you verify, so the check is recorded rather than remembered.
- Do it every cycle, including the ones that feel routine – errors hide in routine cycles.
A team that verifies routinely catches the fabrication before the reader does, which is the entire point.
Common mistakes
- Trusting ground truth. Grounded models still hallucinate; “grounded” is not “verified.”
- Checking with a second model. Agreement between models is not evidence.
- Only checking the prose. Errors hide in figures, causes and labels.
- Softening instead of removing. A doubtful claim should be cut or resolved.
- No source of truth. Without one, verification has nothing to check against.
Frequently asked questions
What is AI hallucination in reporting?
A confident, fluent, false statement – a fabricated figure, a misstated trend, or an invented cause – produced by a model that does not have reliable knowledge of the specific fact.
How often do AI models hallucinate?
It depends on the task. On hard factual questions, rates range from 22% to 94% across leading models. Even on document-grounded summarization, the best models fabricated in about one in seven responses.
How do you catch hallucinated figures in a report?
Verify every figure and cause against a source of truth, separate drafting from checking, flag rather than soften unverifiable text, and log the check for high-stakes reports.
Is a grounded AI model safe for reports?
Safer, not safe. Grounding reduces hallucination but does not remove it, and errors still cluster in specific figures and causes. Verification remains necessary.
Do hallucinations actually cost anything?
They can undermine the report and the function that produced it, because readers treat accuracy as a proxy for reliability. A single fabricated figure invites doubt about every other number.
Which report types carry the most hallucination risk?
Compliance and investor reporting, because the figures are relied on externally. That is where verification should be strictest and where AI-assisted drafts should never reach the reader unchecked.
How do you stop an unverified figure reaching a report?
Require a source for every figure, verify against the system of record, separate drafting from checking, and make approval a gate. Without a source, the figure does not go in.
How do you make verification a habit rather than an effort?
Verify every cycle against the frozen snapshot, check figures before causes, keep a source log as you go, and treat routine cycles with the same rigor as high-stakes ones.
How long does verification take?
With a source log and a checklist, minutes per report rather than hours – and it removes rework that would otherwise take far longer. It is the cheapest insurance in reporting.
What is the most common verification miss?
A cause that was never confirmed. Figures are usually checked; the why often is not, even though it is the part readers act on. Confirm the cause with someone who knows before the report goes out.
Who is accountable if a hallucinated figure reaches a report?
The named approver, which is the point of having one. Accountability stays with a person, not with the tool, and it must be recorded reliably.
Next step
The risk is not AI itself; it is unverified AI. Build a source of truth, verify every figure against it, and never let a draft reach the reader unchecked. See The Human-Verified Reporting Workflow, and have your next report independently checked with a reporting pilot.
Sources
- Stanford HAI, 2026 AI Index and Artificial Analysis AA-Omniscience benchmark: hallucination rates of 22-94% across 26 models.
- Vectara Hallucination Leaderboard (2026): best grounded-summarization fabrication rate at 13.6%.
- OpenAI o3 / o4-mini system card (2025): o3 hallucination on 33% and o4-mini on 48% of PersonQA prompts.
- Stanford RegLab: legal hallucination rates of 69-88% on specific queries.
Numbers are cited from their sources and dated. Where a source is a vendor benchmark, the sample size is stated.