AI When AI gets it wrong: hallucinations, bias and limits
How to fact check an AI answer in five minutes
A repeatable routine for checking what a chatbot told you: which claims to check first, how to verify a citation, and the three questions that catch most invented answers.
The short answer
- Check by consequence rather than by sentence: two or three claims in a long answer carry the decision, and those are the only ones worth five minutes.
- Names, figures, dates, quotations and citations fail far more often than general explanation, and they are also the quickest details to test.
- The core move is going to whoever issues the fact, such as the regulator, the documentation or the statistical agency, rather than to pages that quote each other.
- Asking the chatbot to verify itself proves nothing, because agreeing with what is already in the conversation is the most likely continuation.
- The most common survivable error is a real framework applied to the wrong country, the wrong year or the wrong version, so check which jurisdiction and date an answer assumed.
Do not check everything. Checking every sentence a chatbot produces takes longer than writing the thing yourself, so you would stop doing it by Thursday. Check the claims that carry consequences, in a fixed order, against something that is not the chatbot. Names, numbers, dates and quotations first, then the claim your decision actually rests on. Most invented answers fall apart inside the first two steps, and the whole routine fits in five minutes because you are only ever checking two or three things.
Decide what is worth checking
Two questions sort it: what happens if this is wrong, and how long does confirming it take? High consequence and cheap to check is where all your effort belongs.
Some output barely needs checking at all. When the model is working on material you supplied, summarizing your document, tightening your paragraph, translating your text, restructuring your notes, the facts came from you and the risk sits mainly in distortion rather than invention. Brainstorming does not need verification either, because you are going to evaluate the ideas anyway. Code is a special case: you check it by running it, and a wrong answer usually announces itself.
Everything else is a claim the model supplied from memory, and that pile always needs work. Move it to the front of the queue if any of these apply:
- You will repeat it in public or attach your name to it.
- Money, health, legal rights or safety are involved.
- It is about a specific named person, company or small organization.
- You will act on it once, in a way that is awkward to reverse.
- It concerns something recent, which collides with the limits set out in why an AI does not know about last week.
One thing that carries no information at all is how sure the answer sounds. As the reason a firm tone is not evidence sets out, the register is a style rather than a signal, so it cannot help you triage. Pick by consequence instead.
The claims that fail first
Errors are not evenly distributed across a paragraph. A few kinds of detail fail far more often than the rest, and they happen to be the quickest to test.
| Claim type | What usually goes wrong | Fastest check |
|---|---|---|
| Proper names | Real person attached to the wrong role, company or work | Search the name plus the role and look for an official page |
| Figures and statistics | Plausible magnitude, invented precision, no origin | Search for the issuing body, not the number |
| Dates and sequences | Events placed in the wrong year or wrong order | Check one anchor date on a primary record |
| Direct quotations | The voice is right, the sentence was never said | Exact phrase search in quotation marks |
| Citations | Correct format, nonexistent paper or wrong claim attached | Resolve the identifier, then read the abstract |
| Does this thing exist | A feature, shortcut or rule that was never real | Search the official documentation, not the web at large |
| Calculations | Arithmetic done by prediction rather than computation | Redo it yourself or in a spreadsheet |
The last row deserves its own habit. A model producing a number inside prose is not running a calculation, which is why arithmetic is a known weak spot unless the tool is actually executing code. Any figure that came out of a multiplication, a percentage or a unit conversion should be redone by you, and that takes ten seconds.
The five minute routine
- First thirty seconds: mark the load bearing claims. In a long answer there are usually two or three sentences the rest depends on. Highlight those. Everything else is scaffolding you can ignore.
- To ninety seconds: exact phrase search the most specific string. Take the most distinctive fragment, a full title, a quoted sentence, an unusual pairing of words, and search it inside quotation marks. Zero results for something that is supposed to be documented is close to conclusive. Look past any AI summary at the top of the results page, since that layer can cheerfully describe a thing that does not exist.
- To three minutes: go to whoever issues the fact. The regulator for a rule, the company documentation for a feature, the statistical agency for a number, the court or registry for a case. One hop to the organization that would know beats five hops around sites that are quoting each other.
- To four minutes: read the number in context. Check the year, the units, the denominator and whether it is a measurement or a projection. A real figure attached to the wrong year, the wrong population or the wrong basis is the most common survivor of steps one and two.
- To five minutes: try to disprove it. Search the claim alongside words like disputed or myth, or search the opposite phrasing. If a surprising claim has no critics anywhere, that is itself worth noticing.
Two shortcuts make this faster over time. Keep the primary sources you use often, the tax authority, the regulator, the standards body, the official documentation, in a folder, because step three is most of the work. And check the claim rather than the answer: you are testing one sentence, not grading an essay.
Why asking the chatbot again is not a check
"Are you sure? Please verify" produces a reply, and replies are what you already had too much of. Inside the same conversation the model can see its own previous answer, and continuing agreeably with what is already on the page is the likely move, for the reasons in why a chatbot goes along with you. Confirmation from inside the chat is worth nothing.
Asking in a fresh chat, or a different product, is better but is still only a consistency test. It tells you whether the claim is stable in the training data. It does not tell you whether it is true, and different models trained on overlapping material can be wrong in the same direction. Treat a mismatch as a stop sign and a match as permission to continue checking.
Browsing modes genuinely help, because they give you links. They also invite a new mistake: accepting the citation without opening it. The four step verification in how to check a reference you were given exists because a link that loads is not the same as a source that says what you were told it says.
A worked example
You ask whether you are owed anything for a four hour delay on a domestic flight inside the United States. The answer arrives clean and specific: yes, the airline owes fixed cash compensation scaled by delay length and flight distance, with meals provided after a set number of hours.
It sounds right because it is right somewhere. That is the structure of the European rules on delayed and canceled flights, which do set fixed sums by distance and delay and do require care while you wait. The model has produced a well known framework and attached it to the wrong country, which is a far more common failure than pure invention.
Triage puts this in the check pile immediately: money, a rule, and a claim you would act on. Step two, an exact phrase search of the specific wording, returns European material. Step three settles it: go to the US transport regulator's own consumer pages and the airline's contract of carriage. Federal rules there are built around refunds rather than compensation, so a canceled or significantly changed flight you decline gets your money back, while a delay does not trigger a fixed payment. Anything more comes from the airline's own published commitments, which vary by carrier.
Total time, under two minutes, and the tell was visible before you started: a precise entitlement quoted for a jurisdiction, with no mention of which law created it. Passenger rights, consumer refunds, employment rules and tax thresholds all vary by country and get revised, so this is general information rather than advice, and the regulator's current page is the only version that counts. The underlying reason a plausible framework lands in the wrong place is covered in why a model fills a gap with the nearest matching shape.
What to do when you cannot verify it
Sometimes five minutes is not enough, and the useful skill is recognizing that quickly instead of settling for weak confirmation. Three honest options remain. Mark it unknown and write around it, which is almost always available and almost never chosen. Ask someone with standing: a pharmacist, an accountant, the vendor's support desk, a reference librarian. Or state it with its uncertainty attached, naming what you could not confirm.
Be careful with agreement between sources. Three pages repeating one press release are one source, and search results now contain a growing amount of machine generated filler that will happily echo the claim you are trying to test. How to judge whether a source is worth believing is the companion skill to this routine, and checking a claim before you pass it on covers the version of the problem that arrives through your group chat rather than through a chatbot.
If you keep one habit from all of this, keep the smallest one: before you repeat a name, a number, a date or a quotation that came from an AI, open one page that is not the chat window. That single step catches most of what goes wrong, and it costs less time than the correction would.
Common questions
How do I check an AI answer when I know nothing about the subject?
Check the checkable parts rather than the argument. You do not need expertise to confirm that a named institution exists, that a quoted sentence appears in the document, that a figure matches the agency that published it, or that a cited paper resolves. If those hold up, the answer is at least anchored. If they do not, you can stop without ever judging the substance.
Is it enough to click the links a chatbot gives me?
Only if you read what is behind them. Common failures are a link to a page that never makes the claim, a link to another summary rather than the original, and a real source whose finding is narrower or more hedged than the answer suggested. Skim for the specific sentence that supports the claim, and if you cannot find it, treat the claim as unsupported.
Do AI detectors or verification tools do this for me?
Not reliably. Tools that score text for machine authorship say nothing about whether the content is true, and automated fact checking tools tend to work best on claims that were already widely checked. They can be a first pass on a long document, but the decision about whether a source supports a claim still needs a person reading it.
How often does an AI answer turn out to be wrong?
There is no single honest rate, because it depends entirely on the question. Well documented general topics are usually fine. Obscure, local, recent and numerically precise questions fail far more often, and any published benchmark figure applies to one model, on one test set, at one moment. Judge by question type instead of by an overall percentage.
Should I ask the same question of a second chatbot?
It is a useful cheap test and not a verification. Two independent answers that disagree tell you something is wrong, which is genuinely valuable. Two that agree may simply reflect overlapping training material, including the same bad page on the open web. Agreement moves the claim forward, it does not settle it.