How to check AI output before you rely on it: a manager’s guide to hallucinations and bias.
By Zain M · 2 October 2026 · 11 min read
AI output has to be checked because language models produce fluent text whether or not it is true, and they are wrong most often in exactly the places that matter: specific figures, citations, names, dates, legal and regulatory statements, and anything the model was not given. The five checks that catch most errors take minutes: trace every figure and quote to a source you can open; check every name, date and reference against the record; ask the model to list what it was not given; compare the output against the input document rather than reading it alone; and have a named person sign anything that leaves the firm. The ICO expects accuracy to be a design property of any AI processing personal data, and bias to be monitored where people are scored, and both belong in the workflow, not in a hope.
What a hallucination is, and why it is fluent
A language model produces the most likely next words given what it has seen, which is why its output reads well. It has no separate step that checks whether a statement is true, so when the likely continuation is a plausible-sounding figure, citation or case name that does not exist, that is what it writes, in the same confident tone as everything else. Modern models are far better than the early ones and the business plans of ChatGPT and Claude search the web and cite sources, which reduces the problem, but it does not remove it. The tone gives no signal. The checks have to be structural.
In our own experience running two AI products, the model we trust most is the one that says “the document does not say” rather than filling the gap, and that behaviour is worth more to a professional-services firm than any benchmark. It is also why we design every client system so that a person decides anything that affects a customer.
Where errors cluster
Specific numbers, especially percentages and prices, where a plausible figure is easy to generate. Citations and references, where the format is easy and the existence is not. Names of people, cases, products and organisations. Dates and sequences. Legal and regulatory statements, where the model blends jurisdictions and versions. And anything the model was not given: asked to summarise a contract it has not seen in full, it will summarise what a contract like that usually says. The practical rule is that the more specific and checkable a statement is, the more it must be checked.
| Output type | Risk | The check |
|---|---|---|
| A figure or percentage | High | Open the source it claims; if none, treat as unknown |
| A citation, case or reference | High | Confirm it exists and says what is claimed |
| A name, date or product | Medium | Check against the record you hold |
| A legal or regulatory statement | High | Check the current UK source, not the model |
| A summary of a document you supplied | Medium | Read the output against the input, section by section |
| Prose with no specific claims | Low | Read for tone and sense |
The five checks
Applied to anything that will be sent, filed, published or acted on. They take minutes and they catch most of what matters.
Designing the workflow so the checks happen
Checks that depend on discipline fail on a busy Thursday. Checks built into the workflow do not. Ground the model in your own documents and data so it answers from them rather than from memory, which is what a knowledge assistant on your own material does. Require sources in the output format, so a figure without one stands out. Route anything with a specific claim to a review step and anything with no claims past it. Set an acceptance threshold as a number for any repeated task, measure it on a sample of real inputs before go-live, and keep measuring. And for anything that scores or ranks people, show the reasons with the score so a reviewer can disagree in one click.
Bias, and what the ICO expects
The ICO’s guidance on AI and data protection treats accuracy as a design property of any system processing personal data, and it expects organisations that use AI to make or support decisions about people to understand and mitigate bias. For an SME that means three things. Test the system on a sample that includes the groups it will affect, before go-live. Monitor outcomes by protected characteristic once live, so a skew is found by you rather than by a complaint. And keep a human in the decision wherever the consequence lands on an individual, which is what the Data (Use and Access) Act’s safeguards on significant automated decisions require. Our DPIA template records all three.
Training the team to do this
Most staff have never been shown what a hallucination looks like, and once they have seen three real ones in their own work the checks become habit. A ninety-minute session per team, on that team’s own documents, with the five checks applied to real output, is the most effective training a firm can buy, and it is the part of our training sessions that gets the strongest response.
Common questions
What is an AI hallucination?
Fluent, confident output that is not true: an invented figure, citation, name or statement produced because it was the likely continuation, with no internal check that it is real. The tone gives no signal, so the checks have to be structural.
How do I check AI output for accuracy?
Trace every figure and reference to a source you can open; reconcile names and dates against your records; ask the model what it was not given; read summaries against the source document; and have a named person sign anything that leaves the firm.
Where is AI most likely to be wrong?
Specific figures, citations and references, names, dates, legal and regulatory statements, and anything it was not given. The more specific and checkable a claim, the more it must be checked.
Can AI output be trusted for legal or financial work?
As a draft, yes. As advice, only after a qualified person has checked every specific claim against the current source. Regulators expect that person to be accountable for the result.
How do we reduce hallucinations in a business tool?
Ground the model in your own documents and data, require sources in the output, route claims to review, set an acceptance threshold and measure it on real inputs, and keep a human in any decision that affects a person.
What does the ICO say about AI accuracy and bias?
Accuracy is a design property of any AI processing personal data, and organisations using AI in decisions about people must understand and mitigate bias, test before go-live, and monitor outcomes. The Data (Use and Access) Act keeps the safeguards on significant automated decisions.
Are newer models still wrong?
Less often, and the business plans search and cite, which helps. They are still wrong in the same places, so the checks remain.
Want your team shown this on their own work?
A session per team on their own documents, with the five checks applied to real output and an artefact they keep, from £1,200, delivered within three weeks.