Most AI disappointment in a small business comes from a mismatch of expectations, not from a bad tool. People try AI on a task it does badly, conclude the whole category is overhyped, and stop — or they assume it can be trusted on a task it cannot do, and learn the opposite lesson expensively.
This guide sets the boundary clearly: the tasks AI handles well, the tasks it handles badly, and the verification each one needs. It is the companion to How to Adopt AI in a Small Business.
The useful question is not “can AI do this?” It is “can I check this in minutes?” Where the answer is yes, AI usually helps. Where it is no, it usually misleads.
The three capabilities that matter
Strip away the marketing and AI does three things that are useful to a business.
- It transforms text and structure. Summarize, reformat, restructure, translate, extract, and turn one document into another shape.
- It generates a plausible first draft. Prose, outlines, plans, sequences and code, at a speed no person can match.
- It classifies and extracts at volume. Sort, tag, route, and pull structured fields out of unstructured inputs.
Almost every useful business use case is a combination of these three. A client report is transformation plus drafting. A support workflow is extraction plus drafting. A research briefing is extraction plus synthesis.
The key word in the second capability is plausible. AI produces text that reads as though it were true. That is exactly what makes it useful for drafting — and exactly what makes verification mandatory.
What AI does well
These tasks are where AI reliably earns its place, because the output is checkable quickly.
| Task | Why it works | The check |
|---|---|---|
| Summarizing long documents | Compression is mechanical | Spot-check against source |
| Drafting structured documents | The structure constrains the output | Verify figures and claims |
| Reformatting and extracting | Rules are explicit | Compare to input |
| First-pass research synthesis | Speed of coverage | Open every citation |
| Generating options and variants | Volume of alternatives | Human selects on criteria |
| Turning notes into a narrative | Structure from fragments | Confirm facts |
Notice that every row has a cheap check attached. That is not a coincidence; it is the criterion that makes the list.
What AI does badly
These are the tasks where AI fails, and where the failure is often fluent rather than obvious.
- Arithmetic and totals. Models generate text, including numbers, and can produce internally inconsistent figures. Never let a model perform or restate a calculation that will be relied upon.
- Citations and references. Fabricated sources are a well-documented failure. Reported hallucination rates range widely — roughly 22% to 94% depending on the task and measurement, with one benchmark finding 13.6% of responses grounded (Stanford HAI, AI Index 2026, with Vectara’s leaderboard and OpenAI model documentation). No citation is accepted until a human has opened it.
- Anything requiring your internal context. Without your data and history, output is plausible generalities.
- Recent events. Training cut-offs mean models can be confidently out of date.
- Judgement with consequences. Pricing, hiring, credit and clinical decisions require accountability a model cannot carry.
- Precise legal, tax or regulatory positions. Fluency is not authority. These require a qualified human.
The pattern is consistent: AI fails where correctness cannot be established from the output itself, and where accountability cannot be delegated.
The boundary moves, the discipline does not
Models improve, and tasks that failed last year may work this year. Two things do not change.
- Verification remains mandatory wherever output reaches a client, a filing or a decision. Better models reduce the error rate; they do not eliminate it.
- Accountability remains human. There is no version of a model that can be named as the owner of a client deliverable.
That is why the adoption discipline in this program is built around workflow, data rules and verification rather than around tool selection. The tools will change; the discipline is what makes them safe to use.
Deciding whether a task is suitable
Four questions settle it.
- Is the task frequent? Rare tasks rarely repay the setup.
- Can the output be checked quickly? If not, the use case cannot pay back.
- Does it involve figures or citations? Then verification is non-negotiable, and the workflow must design for it.
- Who is accountable for the result? If the answer is vague, do not proceed.
See The AI Readiness Assessment for the diagnostic version, and Choosing Your First AI Use Case for selection.
How to tell the difference in practice
The boundary between good and bad tasks is easier to see with examples than with principles.
A task AI handles well: turning a long supplier contract into a one-page summary of obligations and dates. The transformation is mechanical, and a person can check the summary against the contract in minutes.
A task AI handles badly: producing the final figures for a management report. The numbers must come from the system of record, and a model that generates them may produce a set that is internally consistent and wrong.
A task AI handles well: drafting a first version of a policy document from a clear brief. The structure constrains the output, and a reviewer edits rather than writes.
A task AI handles badly: confirming whether a claim is legally compliant. The tool can produce a confident answer, and it cannot be the authority.
A task AI handles well: classifying a list of support tickets by topic and urgency. The rule is explicit, and a sample can be checked quickly.
A task AI handles badly: deciding which of two candidates to hire. The judgement carries accountability and consequence, and the check is neither quick nor objective.
The pattern in every pair is the same: the good tasks have a cheap, objective check; the bad ones do not.
A worked example
A professional services firm wants to use AI to speed up its monthly client report. Applying the boundary:
- Keeping AI on the wrong side: asking it to calculate the fee variance and populate the financial table. The figures would look right and might not be.
- Keeping AI on the right side: asking it to draft the narrative commentary from the frozen figures, and to restructure the previous month’s report into this month’s shape.
- The check: the client lead verifies every figure against the finance system and every claim against the delivery record, then signs off.
The result is a report drafted in a fraction of the time, with the numbers produced by the system of record and a named person accountable for the whole. The same task, split across the boundary, works.
Common mistakes
- Assuming fluency means accuracy. A confident, well-written output can be wrong in a single critical detail.
- Using AI for arithmetic. Totals belong in the system of record.
- Trusting citations that were never opened. The most expensive hallucination is a fabricated source.
- Trying AI on a task with no cheap check. The verification cost consumes the benefit.
- Judging the category from one bad task. A mismatch is a selection error, not a technology verdict.
Frequently asked questions
What is AI actually good at in a business?
Transforming and restructuring text, producing plausible first drafts, and classifying or extracting at volume — wherever a person can check the output quickly.
What should I never use AI for?
Arithmetic that will be relied upon, citations nobody has opened, decisions with consequences that require accountability, and precise legal, tax or regulatory positions.
Why does AI make things up?
Models generate plausible text rather than retrieving verified fact, so a fluent answer is not evidence of a correct one. The response is verification, not avoidance.
Can AI be trusted for client-facing documents?
As a drafting mechanism, yes — provided every figure, claim and citation is verified by a named person before delivery. The accountability stays with the human.
Will better models make verification unnecessary?
No. Better models reduce the error rate; they do not remove the need for an accountable human on output that matters.
How do we decide if a task is suitable for AI?
Check four things: frequency, how quickly the output can be verified, whether it involves figures or citations, and who is accountable for the result.
Is AI useful if we only use it for drafting?
Yes — drafting is where most of the value sits in small businesses, because it is the part of the work that consumes senior time. The drafting is fast and the verification is deliberate.
What about AI features built into tools we already use?
They are usually the best starting point, because the data and the workflow already exist inside the tool. The same boundary applies: check what the feature produces before it is relied upon, and keep the skill in your team rather than the vendor’s.
Can AI help with a task we have never done before?
It can produce a plausible first attempt, and a plausible first attempt is not a good one. Novel tasks have no source to check against, so they suit AI as a starting point for a person, not as an output.
Next step
List the tasks you are considering, mark each against the four questions, and start with the one that is frequent and cheap to check. See The AI Readiness Assessment for the full diagnostic, or book an AI adoption call to run it with you.
Sources
- Stanford HAI, AI Index 2026, with Vectara’s hallucination leaderboard and OpenAI model documentation (2025–2026): hallucination rates reported between roughly 22% and 94% depending on task and measurement, with one benchmark finding 13.6% of responses grounded.
Figures are cited from their sources and dated. Where a source is a vendor benchmark, the sample size is stated where published.