Academic Writing

Does ChatGPT Make Up References? What the Fabrication Studies Found

Six studies, four fields, three years. The rate at which ChatGPT invents a reference ranges from single digits to well over half, and the number means nothing without the model version and the date attached to it.

6 min read
Does ChatGPT make up references: a table of studies by model version and year

In June 2025, researchers at Deakin University asked GPT-4o to write six literature reviews on mental health topics. Of the 176 citations it produced, 35 did not exist at all. Another 64 pointed at real papers with the wrong details attached. That is one reference in five invented outright, from a model released more than a year after ChatGPT's citation problem was first measured in print.

Does ChatGPT make up references? Yes, and it still does in 2026, on current models, not only on the version that made headlines in 2023. The open question was never whether this happens. It is how often, and that number moves with the model, the version, the field and the date the test was run.

Does ChatGPT Make Up References?

Yes. A language model predicts the next plausible word given everything before it. It doesn't look anything up unless you add a retrieval or search step in your product. If left to predict what a citation looks like, it will make one that looks correct: author name, year, journal title, volume number, DOI, without verifying any of those things match a real publication. Some of it'll be accurate by accident or by rote memorization. Some of it will be made up from scratch and look exactly the same.

Why a Fabrication Rate Needs a Model Version and a Date

Because the rate is not one number. In the single most-cited study on this, ChatGPT-3.5 fabricated 55 percent of the citations in a set of short literature reviews, and ChatGPT-4, tested by the same researchers on the same 42 topics, fabricated 18 percent. Same study, same prompts, same day of testing. The only thing that changed was the model, and the rate dropped by two-thirds.

Date matters as much as the model name does. GPT-4 in that study meant whatever OpenAI was serving in mid-2023. A 2026 study of ten current models and research agents, checking more than 220,000 citation links in total, found fabrication rates from 3 percent to 13 percent, with the exact figure depending on the specific model and on whether the product used search grounding or a deep research mode. Neither figure is wrong. They describe different things measured almost three years apart.

Field matters too, independent of model. We've also had economics research run through the same GPT-3.5 and GPT-4 pairing find over 30 percent of GPT-3.5 citations invented and a rate still above 20 percent for GPT-4, close to what the multidisciplinary study found, while the medical-systematic-review study, testing what was labeled the March 2023 build of GPT-4 specifically, measured 28.6 percent on a completely different task: pulling references for existing systematic reviews rather than writing new literature summaries from scratch.

What the Published Studies Found

Six studies, run between 2023 and 2026, across medicine, law, economics and general academic writing, converge on the same finding stated six different ways: check every reference an AI tool gives you, regardless of which model produced it or when.

Study and fieldModel and versionDate testedSampleFabrication rate
Walters and Wilder, Scientific Reports (multidisciplinary)ChatGPT-3.5 and ChatGPT-42023636 citations, 42 topics55% (3.5) versus 18% (4)
Buchanan, Hill and Shapoval, The American Economist (economics)GPT-3.5 and GPT-42023Prompts from Journal of Economic Literature topicsOver 30% (3.5), over 20% (4)
Chelli et al, JMIR (medical systematic reviews)GPT-3.5, GPT-4 (gpt-4-32k-0314 build), Bard (PaLM 2.0)Queried July 2023471 references, 11 reviews39.6% (3.5), 28.6% (4), 91.4% (Bard)
Dahl, Magesh, Suzgun and Ho, Journal of Legal Analysis (law)ChatGPT-4 and Llama 22024Random federal court case questions58% (ChatGPT-4), 88% (Llama 2)
Linardon et al, JMIR Mental Health (mental health reviews)GPT-4oTested June 2025176 citations, 6 reviews19.9% fully fabricated, 56% fabricated or flawed
Rao, Wong and Callison-Burch, arXiv preprint (multi-domain, deep research agents)GPT-4.1, GPT-4o-search-preview, Claude 3.5 and 3.7 Sonnet, Gemini 2.52026Over 220,000 citation URLs3% to 13% depending on model

The legal study adds a detail the raw percentage does not capture: the same research found that models struggle to predict their own hallucinations and tend to accept a user's incorrect premise about a case rather than correcting it. A fabricated citation is one failure mode. A model that also cannot flag its own uncertainty is a harder problem, and it shows up most in exactly the field where a wrong citation has the highest cost.

Humanize your own paper

Transform your AI-assisted text and make it sound human, without touching important words or citations.

Get started free

Has ChatGPT Gotten Better Since 2023?

Somewhat, and unevenly. The clearest single before-and-after is in the Walters and Wilder study itself: GPT-3.5 to GPT-4, same day, fabrication cut from 55 percent to 18 percent. Move forward to GPT-4o in mid-2025 and the rate measured by Linardon's team was 19.9 percent, on a harder and narrower task, six literature reviews in a single research field rather than short summaries spread across 42 unrelated topics. Move forward again to the models with search or deep-research features switched on, tested in early 2026, and the range drops to single digits for most of them, 3 to 9 percent, though one deep-research mode still reached 13.3 percent.

None of that is a straight line down. The Linardon number sits between the 2023 GPT-4 figure and the 2023 GPT-3.5 figure, on a model released more than a year after both, because it was tested on a harder task: real literature review generation in a field where a lot of the source material is itself hard to find. A fabrication rate describes a model on a task on a date. Change any one of the three and the number moves, sometimes by a lot.

Model family is not the only variable that still moves the number today. A 2025 study evaluating eight free chatbots, ChatGPT, Claude, Copilot, DeepSeek, Gemini, Grok, Le Chat and Perplexity, on 400 references across five academic fields found that only 26.5 percent of all references generated were fully correct. Grok and DeepSeek produced zero fabricated references in that test. Claude, Copilot and Perplexity were among the worst performers. Two products built on comparable underlying technology, tested the same month, on the same prompts, landed at opposite ends of the range, which is the same lesson as the model-and-date table above, applied to product rather than to version.

Why Fabricated Citations Look Real

A fabricated citation is not an obvious placeholder. Walters and Wilder found their invented entries carried real author names, ones who publish in the actual field, attached to journal titles that genuinely exist, formatted well enough to pass a quick visual check. Linardon's team found something sharper: of the 35 fabricated citations GPT-4o produced, 33 came with a DOI attached, and 64 percent of those DOIs led to a real, published paper on a completely unrelated topic, not a dead link and not an error page, someone else's actual research standing in for a source that does not exist.

How to Check Whether a ChatGPT Citation Is Real

Search the exact title in Google Scholar or the journal's own site rather than trusting the DOI alone, since a working DOI can point at the wrong paper entirely. Confirm the author list, the year and the journal match what actually comes up, not just that something comes up. Do this for every reference a chatbot gives you, not only the ones that look unfamiliar, since the fabricated entries in these studies were built specifically to look ordinary.

Treat an unfamiliar or narrow topic as higher risk than a mainstream one. Linardon's team found fabrication near 6 percent for a well-studied condition, major depressive disorder, and near 30 percent for two less-studied ones in the same set of reviews. The pattern holds outside mental health as well: a model has more real text to draw patterns from in a crowded field, and more room to invent plausibly in a thin one.

An AI citation checker that runs each entry against a real scholarly database automates that search instead of leaving it to a manual lookup for every single reference, which is the part most people skip under deadline pressure, the exact condition under which a fabricated citation is most likely to survive into a final draft.

Finding them in a finished bibliography is its own procedure and has its own page. For the citations that are real, citing ChatGPT in APA is set out step by step.

None of this changes how you format a citation once you have confirmed it is real. That is a separate question: How to cite ChatGPT covers APA, MLA, Chicago and IEEE for the exchange itself, and a citation generator handles the ordinary reference the same way regardless of which style you are writing in. What no citation format can fix is a reference to something that was never published. That check happens before formatting, on every reference, regardless of which model wrote your draft or which version its label claims.

Frequently Asked Questions

Does ChatGPT make up references? Yes, on every model tested so far, in every published study from 2023 through 2026. The rate varies enormously by model version, by field and by date, from roughly 3 percent on the most recent grounded models to well over half on GPT-3.5 in 2023, so a single number is never the full answer.

Sara

Content planner and copywriter at TextPulse. Sara runs the blog day to day, from planning and drafting through to publishing. She writes the practical guides: clear explanations of academic writing problems, aimed at the person who actually has to hand something in.

Stay updated on AI humanization

Get tips on academic writing, AI detection, and humanization delivered to your inbox.

No spam. Unsubscribe anytime.