Academic Writing

Hallucinated Citations: Audit Your Paper Before Submission

A fabricated reference is not a garbled one. It arrives with plausible authors, a fitting title, a real journal and a valid-looking DOI, which is why reading it again will not catch it. Literature-scale audits published in 2026 found the share of papers carrying at least one fabricated reference rose roughly tenfold in three years. Here is the audit to run on your own list first.

Updated on 11 min read
Checklist for a hallucinated citation audit: DOI resolution, title matching, author-year cross-check, and journal name verification

In June 2023, a federal judge in New York sanctioned two attorneys for filing a brief built on at least six court cases that did not exist. Steven Schwartz had asked ChatGPT for supporting precedent, received case names and quoted reasoning that all sounded plausible, and filed the brief without opening a single case himself. Judge P. Kevin Castel fined him five thousand dollars and ordered a letter to every real judge whose name had been attached to a fabricated opinion.

A reference list can end up in the same place. The fix is a hallucinated citation audit run on your own reference list before an editor, a supervisor or a reviewer runs one for you. This is not just a problem caused by drafting with a chatbot. A remembered citation that never quite existed, a co-author's title copied slightly wrong, or a preprint that changed its final citation after acceptance can all leave the same kind of gap in a reference list that a fabricated one does. The audit below catches all of them, regardless of how any single entry got there.

TextPulse AI citation checker verdict list: one reference verified, one matched against Crossref with a corrected citation, and one flagged to verify manually because its DOI does not resolve
Three verdicts on a pasted reference list: verified, match found with a corrected APA 7 copy, and a verify-manually flag for an entry whose DOI resolves nowhere in Crossref, OpenAlex or Semantic Scholar.

How Common Is a Hallucinated Citation, Really?

The range depends heavily on which model produced the draft. Walters and Wilder, publishing in Scientific Reports in 2023, tested reference lists generated by GPT-3.5 and GPT-4 against real academic writing prompts: 636 citations drawn from 84 AI-generated reference lists, 42 for each model. Fifty-five percent of GPT-3.5's citations were entirely fabricated, against 18 percent for GPT-4. Even the real ones were not necessarily clean: 43 percent of GPT-3.5's genuine citations and 24 percent of GPT-4's still carried a substantive error somewhere in the author, title, year or venue.

Legal research on AI-generated citations shows an even higher failure rate, but that work measures a different failure, hallucinated case holdings rather than fabricated bibliographic references, so the figure is not one to import into an academic context. What both fields share is the underlying cause: a model producing a plausible-sounding citation is optimizing for what a citation usually looks like, not for whether one actually exists. TextPulse's own piece on whether ChatGPT makes up references looks specifically at that chatbot; this audit works regardless of which tool, or which co-author's memory, produced the entry in front of you.

What the 2026 Literature-Scale Audits Found

Those model-level rates describe what comes out of a chatbot. A separate question, and the one that matters for anyone submitting to a journal, is how often a fabricated reference survives every check between a draft and a published paper. Two audits released within a day of each other in May 2026 answered it at the scale of the literature, and the answer is that the failure is no longer rare.

An audit published in The Lancet screened references across roughly 2.5 million biomedical papers on PubMed Central Open Access, covering January 2023 to February 2026. Several thousand references pointed to work that does not exist, spread across more than 2,800 published papers. The trend matters more than the total.

PeriodPapers with at least one fabricated referencePer 10,000 papers
20231 in 2,8283.5
20251 in 45821.8
2026, first seven weeks1 in 27736.1

That is roughly a tenfold rise across three years, and the 2026 figure is drawn from seven weeks rather than a full year, so it is an early reading rather than a settled annual rate. What it establishes is that fabricated references are reaching print in peer-reviewed venues, having passed authors, co-authors, editors and reviewers on the way.

The companion study, Zhao and colleagues on non-existent citations, widened the net to 111 million references across arXiv, bioRxiv, SSRN and PubMed Central and put a conservative floor of 146,932 hallucinated citations on 2025 alone. Its more useful finding concerns distribution rather than volume: fabricated references cluster in fields with rapid AI uptake, and among small and early-career author teams. If you are a solo author or a small group without a research office checking your bibliography, you are in the group where this goes wrong most often, which is an argument for running the audit yourself rather than assuming someone downstream will.

Which Models Fabricate Most, and What Actually Reduces It

The honest answer is that no rigorous public benchmark ranks today's major models against each other on citation fabrication specifically. Small informal comparisons circulate, typically a few dozen references run through two or three chatbots by hand, and the samples are too thin to support a ranking between them. Treating one of those as a league table is how a confident and wrong claim gets into a paper.

What does hold up is the generational comparison in the Walters and Wilder data above, where the newer model fabricated roughly a third as often as the older one.

Model testedCitations entirely fabricatedGenuine citations still carrying a substantive error
GPT-3.555%43%
GPT-418%24%

Three patterns replicate across the literature and are worth more than any single ranking. Larger models fabricate less than smaller ones. Models that can actually search and retrieve sources fabricate substantially less than models answering from their parameters alone, which is the single largest effect anyone has measured. And obscure sources are far riskier than well-known ones, because a heavily cited paper appears in training data often enough to be reproduced correctly, while a niche one is more likely to be reconstructed from fragments.

The practical reading is that no model is safe enough to skip the check, and the entries most likely to be wrong are exactly the ones you are least able to verify from memory: the specialist citation, the obscure conference paper, the source you added last.

What an Automated Check Catches, and Where It Stops

Running the list through a checker first is faster than resolving forty DOIs by hand, and this site's free AI citation checker does exactly that against Crossref, OpenAlex and Semantic Scholar. It is worth knowing what that class of tool does badly before you rely on one.

A 2026 evaluation of five hallucinated-citation detectors, Badalova and Mayr, found they give useful early warnings but stay constrained by four things: extracting references cleanly from a document, incomplete metadata in the entry itself, patchy database coverage, and registries that disagree with each other. Only the first of those disappears when you paste a list rather than upload a PDF.

The consequence for your own audit is a rule about interpretation. A checker is strongest at the finding that matters most, a DOI that resolves to a different paper than the entry claims, which is close to conclusive. It is weakest on books, chapters, theses, proceedings, very recent work and non-English venues, all of which are thinly indexed everywhere and routinely come back unfound while being perfectly genuine. Unfound is a prompt to check by hand, never a verdict. The manual steps below are what you apply to whatever the automated pass could not settle.

Humanize your own paper

Transform your AI-assisted text and make it sound human, without touching important words or citations.

Get started free

Resolve Every DOI Before You Submit

A DOI, a Digital Object Identifier, is supposed to resolve to exactly one record through the doi.org resolver. Add the DOI from your reference list to the resolver and a working one lands on the publisher's page for that exact article. A DOI that resolves to an unrelated article, a different journal entirely, or nothing at all is one of the fastest hallucination checks available, because a model inventing a citation usually invents a plausible-looking DOI string to go with it rather than reusing a real one.

A DOI that resolves correctly is not automatic proof the citation is genuine. Some fabricated references borrow the real DOI of an unrelated paper, which resolves without error and points to a real article that has nothing to do with the claim it is supposedly supporting. Read the page the DOI lands on, not just the fact that it loaded.

Match Every Title Against a Real Database

Search the exact title (in quotes) in a database that covers your field, not the open web. Use something like Crossref, PubMed for biomedical work, or your discipline's own index. If you're using Crossref, you can use their free guest search to look up a reference by title and first author, or by author and page number if you aren't sure about the title. This will return the DOI and full metadata for whatever they have on file. If there's no match, or if what comes back is a real paper with an entirely different set of authors, you could be seeing the second kind of hallucination: a broken DOI.

A title search catches a specific failure a DOI check does not: a citation with a working, real DOI attached to the wrong title, because the number and the text were assembled separately and never actually checked against each other. If the database's title does not match your reference list's title closely, treat that as a fabrication until you can explain the mismatch, not as a typo to quietly correct.

Cross-Check the Author List and the Publication Year

A hallucinated citation often gets the author right and the paper wrong, because a real, prominent name in the field is exactly what a model reaching for a plausible citation will produce. Search the author's own publication list, on the database you already have open, for the year your reference list claims. If that author published three papers that year and none of them match your title, the citation is very likely invented, built from a real name and a real year that never actually went together.

The reverse pattern shows up too: a real paper, correctly titled, attached to the wrong year, usually the year of a preprint rather than the year of final publication, or a conference version cited with the journal's later publication year. Neither is fabrication in the audit's strict sense, but both send a reader to a citation manager or a co-author who then cannot find what you cited, close enough to the same problem to fix during the same pass.

Spot a Plausible but Invented Journal Name

An invented journal name is built to sound exactly like a real one: a plausible field, a plausible adjective, a name close enough to an actual journal that nobody stops to check it. "The Journal of Applied Cognitive Studies" is invented; "The Journal of Applied Psychology" and "Journal of Cognitive Science" both exist, and a name gluing pieces of the two together reads as familiar for exactly that reason.

Search the journal name on its own, in quotation marks, and look for a publisher's own journal page and a current ISSN, not just a hit on some other site quoting the same reference. A real journal has a homepage listing an editor, a current volume and issue, and a legitimate ISSN registered to it. A name with no publisher page, no ISSN, and no other article citing work from it is the clearest sign in the whole audit that the reference does not exist.

Turn Your Hallucinated Citation Audit Into a Repeatable Checklist

Run these checks in order, on every reference, before the list is considered final. The order matters less than covering all four; a citation can pass one check and still fail another.

CheckWhat it catches
DOI resolutionA DOI that leads nowhere, or to an unrelated article
Title-to-database matchA real DOI attached to the wrong title, or a title that returns no result at all
Author-plus-year cross-checkA real author's name attached to a paper they did not write that year
Journal name verificationA plausible-sounding journal with no ISSN, no publisher page, and no real existence

TextPulse's AI citation checker runs a version of this pass automatically across a full reference list, faster than working through each of the four checks by hand on a long bibliography, though reading what it flags is still worth doing yourself rather than accepting or rejecting a flag on trust. For the DOI step specifically, a dedicated DOI to citation lookup saves the trip to the resolver for each entry one at a time.

A clean bibliography is one of several checks before a manuscript goes out. Reducing an AI score for journal submission is covered separately, and so is the response to reviewers letter that follows it.

None of this replaces reading your own paper once more before you submit it. If a section still reads as generic after the facts check out, an AI humanizer can smooth the prose without touching a single citation; TextPulse's guide to humanizing AI text in academic writing covers that pass section by section. A reference list that resolves cleanly and a manuscript that reads as your own voice are two separate checks, worth running back to back rather than trusting one to cover the other.

Related research: the findings above are examined at scale in Do AI Models Invent References? A Verification Audit of 1,500 AI-Generated Citations, a TextPulse Research working paper with open data, code and a citable DOI.

Frequently Asked Questions

A hallucinated citation audit is a repeatable check run on a full reference list before submission: resolving every DOI, matching every title against a real database, cross-checking the author list and year, and confirming the journal name actually exists. It catches fabricated references regardless of whether a chatbot, a citation manager error, or a misremembered detail put them there.

Sara

Content planner and copywriter at TextPulse. Sara runs the blog day to day, from planning and drafting through to publishing. She writes the practical guides: clear explanations of academic writing problems, aimed at the person who actually has to hand something in.

Stay updated on AI humanization

Get tips on academic writing, AI detection, and humanization delivered to your inbox.

No spam. Unsubscribe anytime.