Data-to-commentary is the automation that turns figures into narrative – the system that reads a variance and writes the sentence explaining it. It is one of the most attractive uses of AI in reporting, and one of the most dangerous, because it produces exactly the kind of fluent, plausible commentary that is hardest for a busy reader to question.
This guide covers what data-to-commentary is, where it fails, and the controls that make it safe. It is the companion to Using AI to Draft Reports.
Automated commentary is confident by design. The controls exist because confidence is not the same as accuracy.
What data-to-commentary is
Data-to-commentary takes structured data – variances, trends, thresholds – and generates narrative automatically: “Revenue rose 12%, driven by renewals; margin held as costs scaled in line with volume.” The appeal is obvious: the commentary writes itself, and the narrative is never late.
The risk is equally obvious once stated: the system generates the what accurately from the data, but the why is often inferred rather than known. And it is the why that a reader acts on.
Where it goes wrong
The failure modes are specific and repeatable.
- Invented causes. The system attributes a movement to a reason the data does not support.
- Threshold noise. It explains a movement that is too small to interpret.
- Misapplied rules. Logic that flags or explains the wrong segment.
- Missing context. It describes a change without the known reason the team has.
- Over-confidence. Language that asserts certainty the data does not carry.
None of these is a reason to avoid the technology; each is a reason to control it.
The control principle
The safe principle is simple: automate the description, keep the explanation human.
- Description – what moved, by how much, against what – is derivable from the data and safe to automate.
- Explanation – why it moved – is a fact held by a person, not a pattern in a table, and must be confirmed.
Systems that respect this line add real value: they remove the mechanical narration and leave the judgment in place. Systems that cross it produce confident commentary that is sometimes wrong, which is the worst outcome in reporting.
Rules for safe data-to-commentary
If you automate commentary, build in these rules:
- Only describe what the data proves. No causal claims from the numbers alone.
- Suppress explanations below a materiality threshold. If a movement is too small, say so.
- Flag, do not explain, anomalies. “This change is unusual; investigate” beats a fabricated cause.
- Leave a slot for human context. The known reason is added by a person, not generated.
- Require human sign-off on every generated narrative before it is published.
For high-stakes reports, written this way, automated commentary is a drafting aid rather than an unaudited author.
Testing a data-to-commentary system
Before you trust one, test it against known outcomes.
- Back-test on last period’s data, where you know the real reasons.
- Seed known anomalies and check whether the system flags or fabricates.
- Compare generated commentary to the human version for accuracy and omission.
- Track its error rate over cycles, and treat any causal claim as high-risk.
A system that has never been tested against a period you understand is a system you do not yet understand.
When data-to-commentary pays off
Despite the risks, data-to-commentary is worth it in specific conditions.
- High-volume, low-variance reporting: many similar reports where the description is mechanical.
- Routine variance commentary: the “up 4% against plan” sentences that need no judgment.
- Early drafts for a human editor: the system produces the description, the human adds meaning.
- Standardized recurring reports: where the shape repeats and only the numbers change.
It pays off least where every movement is strategic, contested or consequential – board packs in a turbulent quarter, for example, where the why matters most.
Designing the human-context slot
The safest data-to-commentary systems build in a place for human meaning.
- Leave the cause blank rather than generating one.
- Prompt the owner to fill the known reason.
- Treat an unfilled slot as a flag, not as optional.
- Keep the finished narrative incomplete until the context is added.
A system designed this way cannot publish a fabricated cause, because the cause is never generated – only a slot for a human to complete.
Measuring a data-to-commentary system
If you run one, measure it like any other control.
- Description accuracy – does it state the movements correctly? Should be near 100%.
- Causal error rate – how often does it assert a cause? Should be zero if causes are not generated.
- Flag precision – do its anomaly flags correspond to real anomalies?
- Reviewer correction rate – how much human editing is needed per report.
A system whose description is accurate and whose causes are always human-supplied is doing its job. One with any causal error rate is not yet safe to run without close review.
Common mistakes
- Letting the system explain causes. Causes are facts, not patterns.
- No materiality threshold. Every fluctuation gets a confident sentence.
- No human sign-off. Generated commentary reaches the reader unaudited.
- No back-test. The system is trusted before it is tested.
- Treating automation as authorship. The system drafts; a person owns the meaning.
Frequently asked questions
What is data-to-commentary?
The automation that turns structured data into narrative – reading variances and trends and generating the sentences that describe them.
Is automated commentary safe for reports?
Only with controls: describe what the data proves, suppress explanations below a materiality threshold, flag anomalies rather than explaining them, add human context, and require sign-off before publication.
Why is data-to-commentary risky?
Because it generates the why as well as the what, and the why is often inferred rather than known. Fluent commentary with an invented cause is the most dangerous output in reporting.
How do you test a data-to-commentary system?
Back-test it on a period whose real reasons you know, seed known anomalies to see whether it flags or fabricates, compare its output to the human version, and track its error rate over cycles.
Can automated commentary replace human writing?
It can replace the mechanical description. It cannot replace the explanation or the judgment, which remain human – which is why sign-off is required before publication.
When is data-to-commentary worth the effort?
In high-volume, low-variance reporting where the description is mechanical and the shape repeats. It pays off least where every movement is strategic, contested or consequential.
How do you design a safe data-to-commentary system?
Build it to describe only what the data proves, suppress explanations below a materiality threshold, flag anomalies, and leave a slot for human context that must be filled before publication.
How do you measure whether it is working?
Track description accuracy (near 100%), causal error rate (should be zero if causes are not generated), flag precision, and how much human editing each report needs.
Should data-to-commentary run on every report?
No. Run it where the description is mechanical and the variance routine. For contested or strategic reports, the why matters most and should stay human.
What if the system cannot determine a cause?
Leave it blank and flag it. A blank the owner fills is safe; a generated guess is not.
How do you stop a data-to-commentary system over-explaining?
Set a materiality threshold and require the system to stay silent below it. Most noise in automated commentary comes from explaining movements too small to matter.
Is data-to-commentary the same as AI drafting?
No. Drafting produces narrative from a brief; data-to-commentary generates narrative directly from the data. The second is riskier, because the cause is inferred from the numbers rather than confirmed by a person.
How do you choose the materiality threshold?
Set it where a movement becomes meaningful to the reader – typically a percentage of the line, or a level that changes a decision. Movements below it get no narrative, because explaining noise creates doubt. Review the threshold as the business changes, or it slowly stops matching what the business considers material. A stale threshold either explains too little or explains too much, and either way it costs the reader trust in the commentary.
Next step
If you automate commentary, automate the description and keep the explanation human. Test the system against a period you know before you trust it. See The Human-Verified Reporting Workflow for the controls, and book a reporting pilot to run a cycle with us.
Sources
- Stanford HAI, 2026 AI Index and Artificial Analysis AA-Omniscience benchmark: hallucination rates of 22-94% across 26 models, worst on specific figures and causes.
- McKinsey, 2025 State of AI: defined human-validation processes distinguish high performers.
Numbers are cited from their sources and dated. Where a source is a vendor benchmark, the sample size is stated.