In plain terms
A model writes what is likely to come next, and it has no built-in signal that separates “I know this” from “this sounds right”. When it lacks the information, it tends to fill the gap with something plausible: a court case that does not exist, a statistic nobody measured, a product feature you never built. The prose is as polished as when it is correct, which is what makes the errors hard to spot.
Why it matters
Hallucination is the main reason AI output cannot be used unchecked in legal, medical, financial and customer-facing work. It cannot be eliminated, only reduced and managed: by giving the model the source material, asking for citations, permitting it to say “I don't know”, and placing verification where errors are costly. The business question is what a wrong answer costs and who catches it.
Example
A lawyer files a brief drafted with an AI assistant. Opposing counsel cannot find six of the cited cases, because they do not exist; the model generated realistic names, courts and dates. The court sanctions the lawyer. The model did nothing unusual. Nobody had checked.
Most often confused with
Hallucination vs. Error
All hallucinations are errors, but they are a specific kind: fabrication where knowledge is missing. A model that miscalculates a sum or misreads a clause in a document it was given has made an ordinary error. The distinction matters because the remedies differ: grounding and citations for hallucination, clearer instructions and checks for the rest.
Under the hood
Causes: next-token training rewards plausible continuations; evaluation and feedback have tended to reward a confident guess over an admission of ignorance; knowledge is compressed lossily in the weights; and the question may fall after the knowledge cutoff or outside the training data. Rates depend on model, domain and task; they are lowest for summarising supplied text and highest for obscure facts, long lists and exact citations. Mitigations: retrieval-augmented generation, requiring quotes from sources, tool use for calculation and lookup, reasoning models, explicit permission to abstain, and automated groundedness checks. Some researchers prefer the term “confabulation”.