Why AI-Written Annotations Describe Papers That Do Not Exist
A fabricated reference can look exactly like a real one until someone searches for it. The annotation underneath is harder to fake: a fluent paragraph describing a paper's method and findings in detail, for a paper nobody wrote.
A supervisor reads the same sentence shape ten times in a row, opening up an annotated bibliography, reading it aloud, one page after another: This study examines the relationship between X and Y, using a sample of participants to explore Z, finding that further research is needed. All entries read fluently. Two of the twelve sources, searched by title, do not exist.
AI hallucinated citations in bibliographies rarely announce themselves in the reference line. A fabricated author name and a fabricated journal can look exactly like real ones until someone searches for them. What gives the fabrication away first, almost every time, is the annotation sitting underneath it: a paragraph confident enough to describe a paper's method and findings in detail, for a paper that was never written.
How to Spot AI Hallucinated Citations in Bibliographies
Read the annotation before the reference line. A citation by itself, an author, a year, a journal, a DOI are all little structured patterns that a language model has seen thousands of times, and they can look right by chance or by memorization even if there's nothing real behind them. An annotation, on the other hand, must say something about what that specific paper really did, with enough detail that you could either find out if it was true or false. If a model doesn't retrieve or read the source, it won't be able to do this reliably, and the result will show up as genericness long before we ever get around to checking the DOI.
Why the Annotation Gives Away More Than the Citation Does
A fabricated citation is guesswork dressed as a reference: a plausible author, a plausible journal, a DOI-shaped string with the right number of digits. A fabricated annotation is guesswork dressed as comprehension, and comprehension is much harder to fake convincingly across a full paragraph. The model has to commit to a sample, a method and a direction of finding, and every one of those details is a place a real paper can contradict it.
That is also why an annotation is easier for a reader to check than a bare citation. A reference is formatted the same way whether or not the source is real, so its shape alone leaks no information. An annotation makes a specific claim about specific content, so it is either supported by the actual paper or it is not, and a few minutes with the real abstract settles which.
Humanize your own paper
Transform your AI-assisted text and make it sound human, without touching important words or citations.
What a Model Writes When It Has Never Actually Read the Paper
Ask a language model to annotate a source it has only a title and a rough topic for, and it will not refuse or hedge. It predicts the paragraph a paper with that title would probably contain, built from thousands of similar abstracts seen in training, and hands it back in the same confident register it would use for a paper it could quote line by line.
It reads like the average of everything ever written near that topic. An honest summary of a study on sleep and memory will give you the name of the real study participants, the real test used and the real effect size reported. A made-up summary would say the study looked at how well people slept and their memory and found there was a strong link between them. It is generic enough to sit comfortably under a different, unrelated title too.
| What to check | A genuine annotation usually has this | A fabricated one usually has this instead |
|---|---|---|
| Numbers | A specific sample size, effect size or percentage matching the real abstract | No numbers, or a round, unverifiable figure that appears nowhere in the actual source |
| Method | The actual method named: a specific test, design or dataset | A vague gesture at "a study" or "an analysis" with no method named |
| Limitations | One limitation the authors themselves raised, specific to their design | A generic limitation true of almost any paper in the field |
| Reference match | The citation resolves to a real record whose title matches the annotation | The reference does not resolve, or resolves to a different, unrelated paper |
None of these checks require reading the whole paper. They require opening it once, long enough to see whether the specific claims in the annotation actually appear in the abstract.
Why a Fabrication Rate Needs a Model Version and a Date
The scale of the problem is measured, though the measurements do not agree with each other, because they were never measuring the same thing. Testing ChatGPT-3.5 and ChatGPT-4 on the same 42 academic topics on the same day in 2023, researchers found 55 percent of GPT-3.5's citations fully fabricated, against 18 percent for GPT-4. Same prompts, same researchers, same day; only the model changed, and the rate dropped by two-thirds.
A separate study, testing hallucinated case facts rather than bibliographic citations, found ChatGPT-4 wrong 58 percent of the time on verifiable questions about real federal court cases, and Llama 2 wrong 88 percent of the time, in research published in 2024. That figure describes a different task on different models and should not be averaged with the citation-fabrication numbers above into one headline rate. What both studies agree on is the shape of the problem: a percentage without a model version and a date attached describes nothing specific. The wider pattern, across six such studies spanning 2023 to 2026, is its own subject: does ChatGPT make up references covers it in full. What matters for a bibliography specifically stays narrower: whether the annotation gives the fabrication away before anyone checks the DOI.
How to Check Every Entry in an AI-Drafted Annotated Bibliography
Start with the reference line: search the exact title in Google Scholar or the journal's own site, and treat a DOI as unconfirmed until it actually resolves to that title, since a fabricated citation can carry a DOI that leads to someone else's real, unrelated paper. An AI citation checker that resolves each entry against a real scholarly database automates that search instead of leaving it to a manual lookup for every single source.
Then check the annotation against the source itself, not against how convincing it sounds. Open the actual abstract and confirm the sample, the method and the finding the annotation claims are actually in it. A free annotated bibliography generator that grounds every summary in the paper's own abstract, rather than composing one from memory, shows what a genuinely source-grounded annotation looks like next to one that was not.
For the entries that survive the audit, citing ChatGPT in APA and citing ChatGPT in MLA are each set out step by step.
None of this changes how you format an entry once you have confirmed it is real. How to cite ChatGPT itself is a separate question, covered in full for APA, MLA, Chicago and IEEE, and a citation generator formats the ordinary reference the same way regardless of style. What neither one catches is an annotation describing a paper nobody wrote, fluent enough that nobody thought to check.
Related research: the findings above are examined at scale in Do AI Models Invent References? A Verification Audit of 1,500 AI-Generated Citations, a TextPulse Research working paper with open data, code and a citable DOI.
Frequently Asked Questions
Because a citation only has to look right: a plausible author, journal and DOI are enough to pass a glance. An annotation has to be right, since it makes a specific claim about a paper's method and findings that either matches the real source or does not. A model that never read the paper can fake the shape of a reference far more easily than it can fake the substance of a summary.
Content planner and copywriter at TextPulse. Sara runs the blog day to day, from planning and drafting through to publishing. She writes the practical guides: clear explanations of academic writing problems, aimed at the person who actually has to hand something in.