goJumboGPT

AI When AI gets it wrong: hallucinations, bias and limits

Why AI sounds certain when it is wrong

Fluency is not confidence and confidence is not accuracy: what the model is really doing when it states something firmly, and the phrasing that reveals a guess.

8 min read How we write

The short answer

  • An AI sounds certain because a confident register was rewarded during training, not because the system measured how well supported the answer is.
  • Asking a chatbot how confident it is produces another generated sentence, which is why the number lands on round values like 85 or 90 and shifts when you push back.
  • The model does compute internal probabilities over its next words, and those carry some signal, but you cannot see them in a normal chat window.
  • The most useful check anyone can run is asking the same question in a fresh chat: real facts come back stable, invented ones drift on names, dates and figures.
  • Over specific detail, tidy symmetry and sourceless phrases like studies show are the common signs that an answer was generated rather than recalled.

An AI sounds certain because certainty is a writing style, and the model learned that style from text written by people who had already checked their facts. There is no gauge inside the system that rises when it knows something and drops when it is guessing. A sentence it has effectively seen ten thousand times and a sentence it is assembling for the first time arrive in the same steady register, at the same speed.

This matters because you read confidence as evidence, and with people that works. A colleague who says "I am fairly sure it is section 12" is reporting something real about their memory. The same sentence from a chatbot tells you only what a helpful sounding answer looks like. The loop that produces each next word has no step where the system consults what it knows and matches its tone to it.

What calibration means, and why it slips

Calibration is borrowed from forecasting. A weather service is calibrated if, across the days it said there was a 70 percent chance of rain, it rained on roughly 70 percent of them. Calibrated is not the same as accurate: a forecaster who says 50 percent every day is useless and perfectly calibrated in a climate where it rains half the time.

Two separate things get called confidence here. The first is the internal probability the model assigns to each possible next chunk of text. The second is what it says in words when you ask how sure it is. Only the first is a number the system computes. Those internal probabilities do carry information: on questions with a short, checkable answer, a model tends to concentrate probability on one option when it is right and spread it out when it is not. The relationship is loose, it varies by model and subject, and you cannot see it in a normal chat window, but it is there.

What happens next buries the signal. After the raw model is built, it goes through the tuning described in how a model gets turned into a chatbot, where human raters compare candidate answers and pick the better one. Raters prefer answers that are direct, complete and committed. One that opens with "I am not certain, but" reads as weak and scores below one that simply states the thing. The training pressure therefore pushes the confident register everywhere, including onto questions where the support is thin.

There is a real trade off in this. A model tuned to hedge heavily would hedge on the easy questions too, which users find maddening. The cost of avoiding that lands on you, as a tone that no longer varies with how well grounded the answer is.

Why asking how confident are you is not a measurement

When you ask a chatbot how sure it is and it replies "about 85 percent," that number came out the same way every other word did: as a plausible continuation. It is not a readout. Nothing was queried and nothing was counted.

The seams show. Self reported figures cluster on round, socially comfortable values: 80, 85, 90, 95. You will almost never get 37 percent, because that is not what a sentence of that shape contains. The number also moves with your tone. Ask "are you sure about that?" and it drops. Ask "that is right, isn't it?" and it holds. Responding to social pressure rather than to evidence is the same habit explained in why a chatbot folds when you push back.

One version of the question does earn its keep. Absolute percentages are noise, but relative ranking inside a single answer is often useful. "Which claim in what you just wrote is weakest, and why" tends to surface the shakiest sentence, because spotting the obscure detail in a paragraph is a language task the model is good at. Use it to decide where to spend your checking time, not whether to check.

Where a real signal does exist

Some indicators do track accuracy. Mostly they are not the ones people reach for.

SignalWhat it actually reflectsHow much weight it deserves
A stated percentageText that fits the shape of the questionAlmost none on its own
Hedging words: may, typically, I believeWriting style set during tuningLow, though a flat statement about an obscure detail is a flag
Token probabilities from a developer APIHow predictable each word was, not whether the fact is trueModerate for short factual answers
The same answer from three cold retriesWhether the claim is anchored in training rather than invented on the spotStrong as a negative test: disagreement means stop
A source that resolves and actually says itThe evidence itselfThe only thing that settles the question
Answer length, or a long visible reasoning chainEffort spent generating, not effort spent verifyingNone by itself

Token probabilities are the closest thing to an honest uncertainty number, and they are still limited. The figure describes the word, not the world: if two phrasings of the same correct answer are available, they split the probability and both look shaky. The randomness setting covered in the sampling controls behind an answer changes how much variety you see without changing anything the model knows.

The cold retry is the test anyone can run. Ask the same question in a fresh chat, in different words, and compare the specifics rather than the overall shape. Facts that are well represented in training come back stable. Invented ones drift: the date moves a year, the name gains a middle initial, the figure rounds differently. Matching answers are not proof, since a model can be consistently wrong. Conflicting answers are close to proof that you are looking at a guess.

The phrasing that tends to reveal a guess

These are tendencies, not tests. Plenty of confident answers are right. Still, when several show up in one paragraph, check before you rely on it.

  • Over specification. Detail finer than the question needed: a section number you did not ask for, a percentage carried to one decimal place, a precise date for an obscure event. Specificity is cheap to generate and expensive to verify, which is the wrong ratio.
  • Attribution with nobody in it. "Studies show," "experts recommend," "research suggests," with no author, year or title. If you ask for the actual reference, you may end up with the problem described in invented books, cases and studies.
  • Swallowing your premise. You ask how to use a feature that does not exist and get instructions instead of a correction, because a question that assumes something is real makes the confirming answer the likely continuation.
  • No "it depends" where it obviously depends. Consumer rights, tax, benefits and refund windows vary by country and change over time. A single flat answer has quietly picked a jurisdiction for you without saying which.
  • Hedges in the wrong place. Watch where the caution lands. Softening the widely known part while stating the rare part flat is a common inversion, and the pattern behind it is described in why a model fills gaps instead of leaving them empty.

What reasoning modes change, and what they do not

Models that work through intermediate steps before answering, described in what thinking before answering actually means, do improve accuracy on problems with several dependent stages: multi step arithmetic, logic puzzles, code that must satisfy conditions in order. If the task has a checkable structure, more steps means more chances to catch a slip.

What they do not do is verify premises. A chain that starts from an invented fact can be internally flawless and completely wrong, and it reads as more trustworthy than a short answer precisely because you can see the work. Visible thinking demonstrates reasoning. It is not a record of checking, and where the missing fact had to come from the outside world, extra steps do not supply it.

Three tests that take about a minute

  1. Ask cold, twice. The retry test above, run in two fresh chats. Thirty seconds, and it catches a large share of invented specifics.
  2. Ask for a source and open it. Not the title, the page. The failure you are hunting is not only the reference that does not exist, it is the real one that says something narrower than you were told.
  3. Ask what would change the answer. "What would have to be true for this to be wrong?" works well. If a solid counterargument appears instantly, the question was contested and the first answer oversold it. If the model simply abandons its position, it never had an anchor.

How to read a confident answer from now on

Split every answer into the part that came from you and the part that came from the model. Summarizing your document or restructuring your notes carries little risk, because the facts were yours. Anything the model supplied from memory is a claim, and its tone says nothing about its quality.

For those, set one threshold: if you will repeat it in public, act on it with money, or put it in a document with your name on it, it needs a source you opened yourself. The five minute procedure in how to check an AI answer quickly is built for that step. The habit to build is not distrust. It is noticing when you were persuaded by the writing rather than by anything the writing pointed at.

Common questions

Why does the AI give me a percentage if it cannot measure its own confidence?

Because you asked a question, and a percentage is what an answer to that question looks like. The model produces the figure by prediction, in the same way it produces a sentence, so it has the format of a measurement without the substance of one. Use it as a rough ranking within one answer if you like, but never as a probability you can act on.

Does saying are you sure make the answer more accurate?

Sometimes, but not for a reliable reason. The phrase often triggers a second pass in which the model reconsiders and catches a genuine slip. It can just as easily push a correct answer into a needless retraction, because doubt in your message makes doubt in the reply more likely. Treat a changed answer as a signal to check, not as a correction.

Can I see the actual probabilities behind an answer?

Not in a consumer chat app, but developers can. Most model APIs will return the probability attached to each chunk of text the model produced, and low values on a name or a number are a real hint that it was improvising. The catch is that the figure describes word choice rather than truth, so it flags uncertainty without confirming correctness.

Are hedged answers more trustworthy than confident ones?

Not reliably. Hedging is a style the model applies by topic rather than by how well it knows the specific fact, so it often softens general statements while stating rare details flat. What does mean something is a hedge that names what is uncertain and why, such as a rule that varies by country. Vague caution attached to everything means nothing.

Which questions should I assume a confident answer is wrong on?

Anything obscure, local, recent or numerically exact. Small organizations, individual people who are not famous, local legal thresholds, prices, version specific software steps and direct quotations are the usual failure zones. On widely explained concepts the confident tone and the accuracy tend to line up, which is precisely why the exceptions catch people out.