An AI hallucination is an output that a model presents as fact but that has no grounding in its input, source material, or verified training data. The failure is structural, not occasional: large language models such as GPT-4, Claude, Gemini, Llama, and Mistral generate each word by sampling from a probability distribution over a Transformer vocabulary, not by consulting a database of verified statements. That mechanism is what makes a chatbot fluent, and it is also what lets the same model invent a citation, attribute a quote to the wrong author, or recite a confident answer about a court case that never existed. Vendors including OpenAI, Anthropic, Google, and DeepMind describe hallucination as a known property of the architecture rather than a defect a patch will remove (OpenAI, Why Language Models Hallucinate). The piece below walks through what the term means, why next-token prediction produces it, how researchers measure it with benchmarks like TruthfulQA, SimpleQA, HaluEval, and FActScore, and which mitigations such as RAG and RLHF actually move the rate down in production. For wider context on the systems this failure mode lives inside, see how modern generative AI systems work.
What an AI Hallucination Is
An AI hallucination is a model output that is fluent and confident but unsupported by the input provided or by any verifiable fact, a failure distinct from outdated knowledge or incomplete context. OpenAI frames the phenomenon as language models producing statements they cannot back up with evidence from training data or retrieval, a property tied to how the system generates text rather than to a single broken component (OpenAI, Why Language Models Hallucinate). The label covers several behaviors that share one feature: the output lacks grounding. A summary that adds a sentence the source never said is a hallucination. A legal brief drafted in ChatGPT that cites a fabricated case is a hallucination. A code assistant inventing a function that does not exist in the library it claims is a hallucination. A famous early case was Galactica, the science-text model Meta withdrew in 2022 after researchers showed it produced plausible but invented papers and citations. Outdated answers from a knowledge cutoff are a separate failure with its own remedy, and researchers are increasingly careful to keep that line clean. Anthropic's interpretability research underlines the same point from a different angle: a model can produce a confident answer whose internal computation does not match the explanation it offers for that answer (Anthropic, Tracing the Thoughts of a Language Model). The reader-facing test for any AI hallucination is the same in every case: would a competent human checking the source see the claim supported there, or not. For the broader risk landscape this sits inside, see AI risks including bias, hallucination, and regulation challenges.

- Confabulation: plausible-sounding content that contradicts the source the model was given.
- Fabrication: invented facts, citations, names, or statistics with no source at all.
- Knowledge cutoff error: a fact that was accurate at training time but has since changed in the real world.
- Grounded output: a response whose specific claims trace back to retrieved documents or to verified training data the model can point at.
Why Next-Token Prediction Causes AI Hallucinations

An AI hallucination is a predictable consequence of how language models work: they generate each token by sampling from a probability distribution over the vocabulary, not by consulting a verified fact store. OpenAI's own analysis attributes the failure mode to this objective rather than to any single training shortcoming: a system trained to predict the next token will sometimes produce the most statistically plausible continuation even when that continuation is wrong (OpenAI, Why Language Models Hallucinate). The model has no built-in operation called check the fact. It has Transformer weights that encode patterns in training data and a sampler that picks among likely next tokens, with a softmax layer converting raw logits into the probabilities the decoder draws from. When the prompt sits inside a region of the distribution the model saw often, the output tends to be accurate. When the prompt lands in a sparse region, the most probable continuation may simply be a confident guess that pattern-matches similar text the model has seen. GPT-4 producing a fabricated case citation is doing the same statistical operation it does when producing a correct one, and the surface form looks identical to the reader. Decoding choices shape the failure mode too: beam search tends to favor high-likelihood but generic continuations, while temperature sampling can push the model further into invented territory on open-ended prompts. The same mechanism shows up in Claude, Gemini, Llama, and Mistral outputs, which is why DeepMind and Anthropic publish similar disclaimers about factual grounding. Anthropic's tracing research shows that even when a model gives the right answer, its internal pathway can differ from the post-hoc explanation it offers, which is why confidence in the output is not a reliable signal of accuracy (Anthropic, Tracing the Thoughts of a Language Model). The mechanics of token-by-token generation that produce this behavior are covered in how large language models generate text through next-token prediction.
- Tokenize: the prompt is split into tokens and mapped to vectors the language model can process.
- Score: the network computes a probability distribution over the vocabulary for the next token.
- Sample: a token is drawn from that distribution, weighted by temperature and decoding settings.
- Append: the chosen token is concatenated to the input and the loop repeats until a stop condition fires.
- Surface: the resulting sequence is returned to the user with no separate factual accuracy check between generation steps.
Types of AI Hallucination: Confabulation, Fabrication, and Outdated Knowledge

Not every AI hallucination is the same phenomenon: researchers distinguish confabulation (plausible-sounding content that contradicts the source), outright fabrication (invented citations, names, or statistics), and knowledge cutoff errors (facts that were accurate at training time but have since changed). The distinction matters because the mitigation that fixes one rarely fixes the others. Confabulation in a retrieval-augmented setting often responds to better grounding instructions: tighter prompts that direct the model to quote the source verbatim, or schema constraints that force it to cite passage IDs. Fabricated outputs in open-ended generation usually need stricter output constraints or a verification step that re-checks named entities against a trusted index. Knowledge cutoff errors are a retrieval problem, not a generation problem, and the fix is a fresh search rather than a different model. Anthropic's interpretability work shows that confabulation is the harder case because the model's internal pathway can produce an answer that diverges from its own stated reasoning, which means surface-level confidence does not predict whether the output reflects the source (Anthropic, Tracing the Thoughts of a Language Model). Academic surveys of the field draw similar lines, separating intrinsic hallucination (contradicting the input) from extrinsic hallucination (claims that go beyond the input and cannot be verified against it), and labeled benchmarks such as HaluEval and FActScore are built specifically to score these subtypes at scale (arXiv, Survey on Hallucination in Large Language Models).
| Type | What it looks like | Primary driver | Best mitigation lever |
|---|---|---|---|
| Confabulation | Output contradicts the provided source while sounding consistent with it | Weak grounding signal between retrieval and generation | Stricter grounding prompts and quote-with-citation constraints |
| Fabrication | Invented citation, fake statistic, nonexistent function or API | Sparse training data region; model interpolates a plausible-looking artifact | Output constraints, schema validation, named-entity verification |
| Knowledge cutoff error | Confident answer that was true at training time but has changed | Frozen pretraining snapshot with no live source | Retrieval-augmented generation against a fresh index |
How Researchers Measure AI Hallucination Rates
Measuring AI hallucination rates requires purpose-built benchmarks because standard capability tests like MMLU do not isolate fabrication from genuine knowledge gaps. Several evaluations have become reference points. TruthfulQA, developed at Oxford with collaborators including researchers from Stanford, targets common misconceptions, scoring whether a model reproduces a popular falsehood rather than the truth on questions where humans often answer incorrectly. SimpleQA, from OpenAI, targets unambiguous short-answer factual recall, where a single correct answer exists and the model either retrieves it or invents one. HaluEval and FActScore both score model outputs against reference documents to isolate intrinsic and extrinsic errors. Domain-specific suites add another layer: PubMedQA and BioASQ probe biomedical question answering against curated literature, surfacing the kind of confident invention that is most dangerous in clinical contexts. Vectara publishes a public Hallucination Leaderboard that ranks GPT-4, Claude, Gemini, Llama, Mistral, Cohere, and DeepSeek models on summary faithfulness. Other factuality datasets, including Natural Questions, TriviaQA, HotpotQA, FEVER, RAGTruth, and CRAG, probe retrieval grounding and citation accuracy under different question shapes. Summarization-faithfulness suites such as SummEval, XSum, and FRANK extend the same measurement to abstractive summaries. Academic surveys of the hallucination landscape group these benchmarks under the broader effort to separate factual accuracy from general capability, alongside taxonomies that distinguish intrinsic from extrinsic errors (arXiv, Survey on Hallucination in Large Language Models). The benchmarks have limits. TruthfulQA is biased toward English-language misconceptions and can reward evasion as much as truth. SimpleQA depends on the answer index staying fresh; a question with a moving answer is a measurement of retrieval, not of fabrication. Neither benchmark catches confabulation against arbitrary user-supplied sources, which is the failure mode that matters most in retrieval-augmented production systems such as ChatGPT and Perplexity. For the wider picture of what evaluation suites can and cannot tell you about a model, see how AI model evaluation benchmarks work and where they fall short.
| Benchmark | What it measures | Format | Known limit |
|---|---|---|---|
| TruthfulQA | Whether the model reproduces common misconceptions on questions humans often answer wrong | Open-ended and multiple-choice variants | English-centric; can reward evasive non-answers |
| SimpleQA | Factual recall on unambiguous short-answer questions with a single correct answer | Short-form question and answer pairs | Answers can drift as the world changes; tests retrieval as much as recall |
| MMLU (for contrast) | Broad capability across 57 subjects | Multiple choice | Does not isolate hallucination from genuine knowledge gaps |
Mitigation Approaches: RAG, RLHF, and Constrained Outputs
Reducing AI hallucination rates requires layered interventions because no single technique eliminates the failure mode at the architectural level. Retrieval-augmented generation (RAG) is the most widely deployed lever, and it is the pattern that Perplexity and the search-grounded mode of ChatGPT both rely on: the system fetches relevant documents at query time and supplies them to the language model as context, which lowers the rate of confident invention on factual questions. OpenAI's hallucination guardrails documentation treats RAG as a mitigation layer rather than a complete fix, and pairs it with output constraints and verification steps that re-check named entities or numeric claims before the response is returned to the user (OpenAI, Developing Hallucination Guardrails). Reinforcement learning from human feedback (RLHF) is the second lever. The training procedure rewards responses humans rate as accurate and helpful and penalizes confident wrong answers, which shifts the probability distribution toward grounded behavior at inference time. RLHF reduces hallucination frequency but does not remove the failure mode, because the underlying mechanism that produces confident outputs from sparse training regions remains in place. Anthropic's tracing research underscores why the lever is partial: even well-aligned models can produce outputs whose stated reasoning does not match the internal pathway that produced them (Anthropic, Tracing the Thoughts of a Language Model). Google's developer documentation describes a third layer for the Gemini API: grounding constraints, safety filters, and structured output formats that bound what the model is allowed to produce on a given call (Google AI for Developers, Gemini API Safety Guidance). For the wider training picture that RLHF sits inside, see how supervised training and reinforcement learning shape model behavior.
- Retrieval-augmented generation (RAG): fetch relevant documents at query time and supply them as context so the model has source material to ground its response.
- Reinforcement learning from human feedback (RLHF): tune the model so it prefers accurate, grounded responses over confidently wrong ones, shifting the probability distribution at inference.
- Constrained output formats: require structured JSON, schema-conforming fields, or verbatim quote constraints that narrow the space the model can generate into.
- Verification steps: re-check named entities, citations, or numeric claims against a trusted index before returning the response.
- Refusal calibration: train the model to say it does not know when retrieval returns nothing relevant rather than filling the gap with a fabricated answer.
When AI Hallucination Risk Is Highest
AI hallucination risk is not uniform across use cases: tasks that require precise factual recall, citation accuracy, or medical or legal reasoning carry substantially higher exposure than open-ended creative generation. A model writing a marketing tagline can afford to invent flourishes that a model drafting a contract clause cannot. The difference comes down to verifiability and consequence. When a wrong answer carries operational or legal weight, the architectural property that produces confident plausible text is no longer benign, and the mitigation stack has to do more work. Biomedical question-answering benchmarks such as PubMedQA and BioASQ exist precisely because the cost of a fabricated dosage figure or invented study is asymmetric, and clinical deployments need a measurable floor before they ship. OpenAI's guardrails documentation flags the same pattern, advising tighter retrieval grounding and verification steps for production deployments where factual accuracy matters (OpenAI, Developing Hallucination Guardrails). Google's Gemini API guidance points to grounding constraints and safety filters as the lever for high-stakes deployment paths (Google AI for Developers, Gemini API Safety Guidance). For an overview of the broader regulatory and risk picture this sits inside, see AI risks including bias, hallucination, and regulation challenges.
- Legal and medical reasoning: a fabricated citation or dosage figure carries real downstream harm; production systems built on GPT-4 or Claude in these domains require human review.
- Citation generation: any task that names papers, cases, or authors is a known high-risk surface because plausible-sounding fake references are easy for the model to produce.
- Code suggestions for unfamiliar libraries: assistants can invent functions or arguments that do not exist in the library version the developer is using.
- Numeric and statistical claims: models can produce confident figures that look authoritative but trace to no source in retrieval or training data.
- Open-ended summarization of long documents: extrinsic additions creep in when context exceeds what the model can attend to reliably, producing claims that go beyond the source.
What Developers Can Do To Reduce AI Hallucination Today
Developers shipping AI hallucination mitigations work within a consistent set of techniques that reduce exposure without waiting for the next model generation. The mechanism-first view explains why each lever helps: every technique either narrows the probability distribution the model samples from, supplies retrieved grounding the generator can quote, or adds a verification step after generation. OpenAI's hallucination guardrails cookbook lays out the practical pattern: constrained outputs, retrieval grounding, and post-generation checks combined as defense in depth rather than treated as alternatives (OpenAI, Developing Hallucination Guardrails). Google's Gemini API guidance recommends the same layered approach, with grounding and safety filters configured per call (Google AI for Developers, Gemini API Safety Guidance). The net effect is a lower hallucination rate without a model change, and a system that fails closed rather than confidently producing a fabricated answer when retrieval comes back empty.
- Wire retrieval before generation: supply the model with documents at query time so the response has grounded source material to quote.
- Constrain the output schema: require structured JSON or fixed fields so the model cannot drift into free-form invention.
- Use prompt engineering to demand citation: ask the model to quote the source verbatim and label any claim it cannot back up with a passage ID.
- Verify named entities and numbers: run a second pass that re-checks people, places, citations, and figures against a trusted index.
- Calibrate refusal behavior: design the system to return I do not know when retrieval is empty rather than letting the language model fill the gap.
- Log and sample for review: capture hallucination candidates in production traffic and feed them back into evaluation suites such as TruthfulQA-style probes tailored to your domain.
References
- OpenAI, Why Language Models Hallucinate
- Anthropic, Tracing the Thoughts of a Language Model
- Google AI for Developers, Gemini API Safety Guidance
- OpenAI, Developing Hallucination Guardrails
- arXiv, Survey on Hallucination in Large Language Models
Further reading
Frequently Asked Questions
Does retrieval-augmented generation eliminate AI hallucinations?
RAG reduces AI hallucination rates by grounding responses in retrieved documents, but it does not eliminate them. A model can still misread, misquote, or fail to retrieve the relevant passage, producing a hallucination even when a correct source exists in the index. OpenAI's hallucination guardrails documentation treats RAG as a mitigation layer, not a complete fix.
What is the difference between a hallucination and a confabulation?
Confabulation is a specific subtype of AI hallucination where the model generates plausible-sounding content that contradicts the source it was given. Fabrication, by contrast, refers to invented facts with no source at all. The distinction matters for choosing the right mitigation: confabulation responds to better grounding prompts, fabrication requires stricter output constraints or citation verification steps.
Which benchmarks measure AI hallucination rates?
TruthfulQA and SimpleQA are the two benchmarks most widely used to measure AI hallucination rates across model families. TruthfulQA tests whether models reproduce common misconceptions; SimpleQA measures factual recall accuracy on unambiguous short-answer questions. Standard benchmarks like MMLU do not isolate hallucination from genuine knowledge gaps.
Does fine-tuning or RLHF stop a model from hallucinating?
RLHF reduces AI hallucination frequency by training models to prefer accurate, grounded responses over confidently wrong ones, but it does not remove the failure mode. Anthropic research on tracing model thoughts shows that even well-aligned models can produce outputs that diverge from their internal processing, meaning hallucinations persist at lower rates across all current production models.









