The most common serious AI incident in a small business is not a data breach in the conventional sense. It is a confidential document pasted into a consumer AI tool by someone trying to do their job well. The tool worked, the answer was useful, and the business’s data now sits in a service nobody assessed.

This guide covers what to prohibit, how to assess a tool’s data handling, and the controls that make the rule practical. It is part of the Governance, Risk & Data pillar.

The rule is not “never use AI”. It is “never put this category of data into that tool”.


Why this is the control that matters most

Two findings frame the problem. In EOG’s synthesis of 2025–2026 research, a study by Harmonic Security (2025) found 26.4% of employees had pasted confidential company data into a generative AI tool — meaning the behavior is already common in businesses that have not yet addressed it. Separately, IBM’s Cost of a Data Breach Report 2026 put the global average cost of a breach at a record $4.99 million, with organizations under 500 employees averaging $3.31 million.

The first finding is about behavior and the second is about consequence. Together they explain why the data rule is the highest-priority control in an AI policy, and why it needs to be practical rather than absolute.


What never goes into an unapproved tool

A list, not a principle.

  • Client confidential information and anything covered by a non-disclosure agreement.
  • Personal data, especially employee records and any special-category data.
  • Regulated records — financial, health, legal — where handling is prescribed.
  • Credentials and secrets — passwords, keys, access tokens.
  • Anything under a contractual restriction on processing or sub-processing.
  • Source material for a client deliverable, where the client owns it.

And a default: where the category is uncertain, treat it as prohibited and ask the person who owns the policy.

Two clarifications are worth making, because they cause most of the confusion. First, the rule applies to what you paste, so a summary containing confidential facts is as prohibited as the original. Second, it applies to interfaces as well as chats — many tools ingest documents through file uploads and integrations, where the same rule holds.


The tool-side questions

Whether a category may go into a tool depends on the tool. Five questions settle it.

Question Why it matters
Is input retained, and for how long? Retention determines exposure
Is input used to train models? Training use is a disclosure risk
Where is the data processed? Jurisdiction determines obligations
Who can access it on the vendor’s side? Support and engineering access is a real risk
Can you export and delete it? Determines your exit and your ability to comply with a deletion request

The answers belong in writing, in a due-diligence record, and they should be re-checked on material change — a vendor’s terms can move. See AI Tool and Vendor Due Diligence.


Making the rule practical

A rule that blocks work gets ignored. Three design choices keep it usable.

  • Provide an approved tool that covers the common cases, so staff have somewhere to go.
  • Name a person to ask, in the policy, so uncertainty has a route.
  • Give a workaround for the blocked cases — an approved alternative, a redaction step, or a manual process — so the task can still be completed.

The failure mode to avoid is a policy that prohibits the tool people need for the task they have, without offering a path. That produces the behavior the policy was written to prevent, conducted quietly.


The controls around the rule

Four controls make the data rule hold.

  • An approved-tool list, maintained, with named tools rather than categories.
  • A short data-handling induction, delivered with the policy, covering the categories and the default.
  • Technical controls where available — an enterprise tier with training disabled, retention off, and access controls.
  • A log of what was approved and why, so decisions can be revisited rather than re-litigated.

For most small businesses, the most valuable of the four is the enterprise tier on the main assistant: it converts a policy statement about retention into a contractual one.


When something has been pasted

The response should be defined before it is needed.

  • Establish the facts — what was entered, into which tool, and when.
  • Assess the exposure — which category, whose data, and what the vendor’s terms say about retention and training.
  • Act proportionately — from deleting the conversation, to notifying a client, to following a breach process where the data is regulated.
  • Record it, and fix the cause — usually a missing approved alternative rather than a lapse in care.

This is the same five-step shape as any AI incident, and it is covered in When AI Goes Wrong.


A worked classification

Three everyday examples show where the line falls.

  • A marketing blog post in draft. Internal draft content, no client data, no personal data — generally acceptable in an approved tool.
  • A client’s contract, to be summarized. Client confidential, and the client owns the document. Prohibited unless the tool is approved for that category and the contract permits it.
  • A pricing model for a new service. Internal, non-confidential, but commercially sensitive. Acceptable in an approved enterprise-tier tool; not in a consumer tool, and not where the vendor trains on inputs.

The third case is the one that causes debate, because the data is yours rather than a client’s. The test is not ownership but exposure: what happens to commercially sensitive material if it is retained or used for training. Where the answer is unclear, the default applies.

Common mistakes

  • A principle instead of a list. “Do not share sensitive data” is not actionable.
  • No approved alternative. The rule is ignored, quietly.
  • Assuming the default consumer tier is safe. Retention and training terms differ between tiers.
  • Forgetting uploads and integrations. The rule applies to files, not just chat.
  • No records of due diligence. The position cannot be evidenced when a client, insurer or regulator asks what was assessed and when.
  • No response plan. The first incident becomes an improvised one.

Frequently asked questions

What data should never be entered into an AI tool?

Client confidential information, personal and employee data, regulated records, credentials, and anything under a contractual processing restriction — unless the tool has been approved for that category.

Are AI tools safe for confidential business data?

Enterprise tiers with training disabled, defined retention and access controls are generally acceptable for many categories. Free and consumer tiers frequently retain input and may use it for training, which makes them unsuitable for confidential data.

How do we check a tool’s data handling?

Ask five questions: retention, training use, processing jurisdiction, vendor-side access, and export and deletion. Get the answers in writing and record the assessment.

What if staff have already pasted confidential data?

Establish the facts, assess the exposure against the vendor’s terms, act proportionately — from deleting the conversation to notifying affected parties — and record it. Then fix the cause, which is usually the absence of an approved alternative.

Does the data rule apply to file uploads?

Yes. A document uploaded into a tool is subject to the same rule as text pasted into a chat, and integrations carry the same exposure.

How do we make the rule easy to follow?

Provide an approved tool for the common cases, name a person to ask, and give a workaround for blocked cases so the underlying task can still be done.

Is it acceptable to redact confidential details before using a tool?

Often, yes — provided the redaction removes the confidential content rather than obscuring it, and the summary produced is treated as confidential too. Redaction is a workaround, not a license, and it should be documented in the policy.

What about tools built into software we already use?

The same questions apply. An AI feature inside an existing system is still a data-processing decision, and the vendor terms should be checked for retention and training before confidential material is used.

What if an employee has already used an unapproved tool?

Treat it as a near miss rather than a disciplinary matter: establish what was entered, assess the terms, act proportionately and record it. The purpose is to close the gap, not to attribute blame.


Next step

Write the prohibited-categories list, name the approved tools, and check the retention and training terms of the one you use most. See AI Tool and Vendor Due Diligence and the AI Policy Template, or book an AI adoption call to run the assessment with you.


Sources

  • Harmonic Security (2025): 26.4% of employees have pasted confidential company data into a generative AI tool.
  • IBM, Cost of a Data Breach Report 2026 (July 2026): global average breach cost $4.99 million (up 12%); organizations under 500 employees average $3.31 million.

Figures are cited from their sources and dated. This article is general information, not legal advice; confirm your specific obligations with qualified counsel.