Ask ChatGPT for five peer-reviewed sources on almost any topic, and within seconds you'll have a clean, properly formatted reference list — author names, publication years, journal titles, volume and issue numbers, sometimes even a DOI. It looks exactly like a real citation because it was built from the same patterns real citations follow. The problem is that a meaningful share of what comes back is not real: papers that were never written, real researchers credited with studies they never published, and genuine papers cited with the wrong year, the wrong journal, or the wrong page numbers.
This isn't a rare glitch that a future update will quietly fix. It's a well-documented, peer-reviewed behavior called hallucination, and it shows up across every major AI assistant, not just ChatGPT. Multiple academic studies have measured it directly, and the legal profession has spent the past three years learning the same lesson in open court, with sanctions climbing every year since the first widely publicized case.
This guide explains what citation hallucination actually is, why it happens even in the newest models, what has happened to people who trusted it blindly, and a concrete, repeatable process for verifying every AI-suggested reference before it goes anywhere near your bibliography. If you're instead looking for how to correctly cite ChatGPT itself as a source you used, see our guide on how to cite ChatGPT and AI tools — this article is about the opposite problem: AI inventing sources for you to cite.
What Is Citation Hallucination?
In AI research, "hallucination" describes output that sounds fluent and confident but isn't grounded in anything real — a fact, quote, or source the model presents as true even though it doesn't correspond to reality. Citation hallucination is a specific, especially consequential version of this. The model produces something with every visual hallmark of a genuine academic reference — an author, a publication year, a journal name, a volume and page range, sometimes a perfectly formatted DOI — that either doesn't exist at all, or exists in a meaningfully different form from what's presented.
This isn't unique to one company's technology. It shows up in ChatGPT, Google's Gemini, Microsoft Copilot, and other assistants whenever they're asked to supply sources from memory instead of from a live, verified lookup. Some researchers now prefer the term "confabulation" to "hallucination," since it more precisely describes filling a gap with a plausible-sounding invented answer rather than a random error.
Why AI Tools Invent Citations
Large language models are trained to predict the most statistically likely next word given everything that came before it — not to retrieve verified facts from a database. When you ask for "five peer-reviewed sources on X," the model generates text that matches the structure of a citation it has seen thousands of times during training (author, year, title, journal, volume, pages, DOI), without that structure necessarily being linked to one specific, verifiable, real-world record.
For extremely well-known material — foundational textbooks, landmark studies that get cited constantly — the model has seen the same citation repeated so often that it reproduces it accurately. For a narrower or more specialized question, there may not be enough repeated exposure in the training data, so the model fills the gap with something statistically plausible but disconnected from reality. This explains why hallucination rates vary so widely by topic: research on GPT-4o found fabrication as low as 6% of citations for a heavily studied condition like major depressive disorder, but climbing to 28-29% for less-studied conditions like binge eating disorder and body dysmorphic disorder.
Web-connected AI tools that can search and browse in real time reduce this problem, because they can pull results from an actual search rather than reconstructing a citation purely from memory. But reducing the problem isn't eliminating it: a model can still misread a search result, blend details from two different sources into one reference, or fall back on its own memory when a search comes up empty — and it does all of this with the same fluent confidence whether the citation is real or invented. DOIs are a particularly easy target for confident fabrication, since they follow a simple, predictable format (10.xxxx/xxxxx) that a model can generate convincingly without it being tied to any registry it has actually checked.
Three Ways AI Citations Go Wrong
Not every hallucinated citation fails the same way. Recognizing which pattern you're looking at helps you know exactly what to check.
1. Fully Fabricated
The paper, the authors, and the journal do not exist anywhere. No search will ever turn it up because there is nothing to find.
Illustrative example — plausible on every level, invented on every level.
2. Real Source, Wrong Details
The paper genuinely exists, but the model reports the wrong year, the wrong journal, an incorrect volume or page range, or attributes it to the wrong set of authors. A real 2019 study gets cited as 2021; a paper published in one journal gets attributed to a different, similarly named one. These are harder to catch than a full fabrication because a search for the topic or author will still return something real — just not the exact record being cited.
3. Frankenstein Citations
A real, well-published author is credited with a paper they never wrote, or two genuine papers are merged into a single reference that matches neither one. One published analysis of AI-generated psychology citations found that even references with a technically valid, resolving DOI sometimes pointed to an article whose actual title had nothing to do with what the model claimed — meaning a working link is necessary, but not sufficient, proof that a citation is correct.
Real-World Consequences
In Legal Practice
The case that put AI citation hallucination on the map was Mata v. Avianca, Inc., filed in the U.S. District Court for the Southern District of New York. Roberto Mata sued the airline over an injury from a serving cart, and his attorneys used ChatGPT to help research and draft a brief opposing Avianca's motion to dismiss. The brief cited six court cases — including one widely referenced example, Varghese v. China Southern Airlines — that did not exist. When Avianca's lawyers and then the court itself couldn't locate them, presiding Judge P. Kevin Castel ordered the cases produced. The attorney had, at one point, asked ChatGPT to confirm the cases were real; the tool assured him they were.
In June 2023, Judge Castel fined the attorneys and their firm $5,000 and ordered them to send copies of the sanctions decision to every judge falsely named in the fabricated opinions. It was widely reported as the first major instance of its kind — but it was nowhere near the last. A database maintained by a research fellow at HEC Paris tracking AI hallucinations in court filings has since logged more than 1,200 court proceedings worldwide, with close to 500 attorneys sanctioned across more than 100 countries. Fines have escalated well past the original $5,000, with individual 2025 sanctions running into the tens of thousands of dollars. In one 2026 case, a U.S. federal appeals court sanctioned attorneys who had cited more than two dozen fabricated authorities in a single appeal, and courts have continued sanctioning repeat offenders — lawyers caught citing fake AI-generated cases again after already being warned.
In Academic Publishing
Journals and publishers have faced a version of the same problem at scale. In 2024, Wiley disclosed that AI-assisted "paper mills" had mass-produced manuscripts containing fabricated data, plagiarized text, and hallucinated citations, forcing the publisher to shut down 19 journals and disclose $35-40 million in lost annual revenue tied to the scandal. Separately, one academic publisher retracted more than 100 articles that had been generated or substantially written with large language models without disclosure, with a large share traced back to a single institution.
In response, major academic publishers — including Elsevier, Springer Nature, Wiley, Taylor & Francis, and SAGE — now explicitly require authors to disclose any generative AI use in preparing a manuscript. Several journals have stated plainly that they will reject, withdraw, or retract any submission found to contain fabricated references, regardless of whether that's discovered before or after publication.
How to Spot a Fake or Hallucinated Citation
Before you run a full verification pass, these signs are worth checking first — any one of them is reason enough to treat a citation as unverified.
- The DOI doesn't resolve when you enter it at doi.org, or it resolves to a completely different article
- The exact title returns no matching results on Google Scholar, CrossRef, or PubMed
- The journal name is close to a real one but not quite right — a common sign of a blended or invented reference
- The claimed volume, issue, or page numbers don't match that journal's actual publication history for the stated year
- The citation supports your argument a little too perfectly, with no caveat or nuance — real research rarely lines up this conveniently
- The named author has no other findable publication history, or nothing else they've published relates to the claimed topic
Step-by-Step: How to Verify Every AI-Suggested Citation
- 1
Pull every AI-suggested reference into its own list
Do this before you write another word of your paper around it.
- 2
Search the exact title in Google Scholar
If nothing matches word-for-word, treat the citation as unverified.
- 3
Resolve the DOI through doi.org or CrossRef
Confirm the resolved page shows the same title, authors, and year the AI gave you — not just that a DOI happens to resolve to something. See our full guide to verifying a DOI for the complete process.
- 4
Cross-check PubMed for medical and health topics
PubMed indexes far more rigorously than a general web search for biomedical literature.
- 5
Check the publisher's own site
Confirm the article actually appears in the specific volume and issue cited.
- 6
If you can't verify it, don't use it
Either replace it with a source you can confirm is real, or remove the claim it was supporting.
- 7
Run a final automated check on the whole document
Upload your finished reference list to a reference checker as a last safety net — it catches mismatches across a long document that a source-by-source manual pass can still miss.
How to Use AI Safely for Research
Ask AI to work from a source, not from memory
Summarizing a PDF you've uploaded directly is a fundamentally different — and far more reliable — task than asking an AI tool to "find three studies about X" from what it remembers. Use AI to brainstorm search terms, explain a concept, or tighten your writing, not to supply sources from its own training data.
Never let a model verify its own output
Asking the same AI tool to "confirm this citation is real" is not independent verification — it is exactly what happened in Mata v. Avianca, when ChatGPT reassured the attorney that its own invented cases were genuine. Verification has to happen against an outside source: CrossRef, PubMed, Google Scholar, or the publisher itself.
Build verification into your drafting process
Check each AI-suggested source the moment you add it, rather than saving verification for a single pass at the end. This prevents a citation you never actually checked from quietly becoming "something you already verified" in your own memory by the time you submit.
Frequently Asked Questions
Does ChatGPT know when it's making up a citation?
No. The model has no internal mechanism for flagging its own fabricated content as uncertain — it generates the most statistically likely continuation regardless of whether that continuation happens to correspond to something real. This is exactly why asking the same model to "confirm" its own citation doesn't work: in the Mata v. Avianca case, ChatGPT reassured the attorney that its invented cases were genuine.
Do newer models like GPT-4 or GPT-4o still hallucinate citations?
Yes, at a lower rate than earlier versions but not close to zero. A peer-reviewed comparison found fabrication dropped from 55% of citations under GPT-3.5 to 18% under GPT-4 — real progress. But a 2025 study of the newer GPT-4o model still found fabrication rates around 20% overall, climbing toward 30% for less-studied topics, with a large share of the "real" citations it produced still containing factual errors.
Can I actually get in trouble for submitting an AI-fabricated citation?
Yes, and the consequences have escalated quickly. In academic settings this can be treated as an integrity violation and has led to paper retractions and journal shutdowns. In legal practice, courts have sanctioned hundreds of attorneys worldwide since 2023, with fines climbing from the original $5,000 in the first major case to tens of thousands of dollars in more recent ones.
Is it fine to use ChatGPT to help write my paper at all?
Most institutions and journals now permit AI assistance for brainstorming, outlining, and improving clarity, but require you to disclose that use and to independently verify every factual claim and citation yourself. Always check your specific institution's or target journal's current AI policy, since these are changing quickly.
What's the fastest way to check an entire AI-generated reference list at once?
Upload the full document to a dedicated reference checker rather than verifying each citation one at a time by hand. A reference checker cross-references every entry against sources like CrossRef and flags fabricated or mismatched citations across the whole list in a single pass.
Conclusion
AI tools are genuinely useful for drafting, brainstorming, and organizing your thinking — but they were never built to be a source of verified facts, and citations are exactly the kind of specific, checkable detail they handle worst. The gap between how confident a hallucinated citation sounds and how real it actually is won't close on its own; closing it is now a required step in anyone's writing process, not an optional extra for the unusually careful.
The good news is that verification doesn't have to be slow. A handful of DOI lookups and title searches per source, done as you write rather than saved for the end, catches the overwhelming majority of fabricated or inaccurate AI-generated citations before they ever reach a reader, a reviewer, or a judge.
Check Your AI-Generated Citations Now
Upload your document and catch fabricated, mismatched, or inaccurate references automatically — before your instructor, editor, or a judge does.
Verify My ReferencesRelated Articles
How to Check If a Reference Is Real or Fake
Step-by-step guide to verifying citation authenticity
Read moreHow to Verify a DOI: Step-by-Step Guide
Using doi.org, CrossRef, and PubMed to confirm any DOI
Read moreHow to Cite ChatGPT and AI Tools
The correct citation format for AI tools you used directly
Read moreCitation Checker vs Plagiarism Checker
Why these tools solve different problems
Read more