← The Signal Report Work with me

Report 057 · AI in the Lab

Fake citations in real papers

An automated audit of 2.5 million biomedical papers found that the rate of references pointing at studies which do not exist has risen more than twelvefold since 2023. That is a real finding and it is worth knowing. It is also, I think, the least dangerous version of the problem, because a citation that does not exist is the one kind a machine can catch.

Every report on this site is built on a rule I set for myself before I published anything: I open every source I cite. Not the abstract, not the press release, not somebody's summary of it. The actual document. It is slow, it kills stories I liked, and it is the only part of the method I would defend without qualification.

That rule started as a discipline carried over from the lab, where a number you did not personally watch come off the instrument is a rumor. It has turned out to be a good rule for a different reason, which is the subject of this report.

What the audit did

In May 2026 a team led by Maxim Topaz at Columbia University's School of Nursing and Data Science Institute published a short correspondence in The Lancet titled "Fabricated citations: an audit across 2·5 million biomedical papers." The system behind it is called CITADEL, for Citation Integrity Testing and Detection of Erroneous Literature.

The design is simple, which is what makes it convincing. They took the PubMed Central Open Access corpus from 1 January 2023 through 18 February 2026: about 2.5 million papers containing roughly 125.6 million references, of which 97.1 million were successfully verified. Language models were used to compare each reference's title against the DOI or PubMed identifier attached to it, which is mostly a job of telling a genuine mismatch apart from a formatting quirk like an abbreviated journal title. Anything that still looked wrong was then checked against outside databases. The project describes the standard plainly: every reference was verified against four independent databases, PubMed, Crossref, OpenAlex and Google Scholar, and a reference was classified as fabricated if its title did not appear in any of them.

The result: 4,046 fabricated references across 2,810 papers. Expressed as a rate, the quarterly figure per 10,000 papers indexed in PubMed Central rose from about four in 2023 to 56.9 in early 2026, a more than twelvefold increase. Put the other way round, roughly one paper in 2,828 carried a fabricated reference in 2023, one in 458 in 2025, and one in 277 in the first seven weeks of 2026. The steepest part of the climb began in mid-2024.

One paper in the set contained 18 fabricated references out of 30.

Read the denominator before you panic

Here is where I want to slow down, because two opposite misreadings are available and both are wrong.

The first misreading is that the literature is now mostly fake. It is not. Four thousand bad references out of 125.6 million is a vanishingly small fraction, and 2,810 affected papers out of 2.5 million means better than 99.6% of papers in this corpus had no fabricated reference at all. If you came here for the collapse of science, this is not it.

The second misreading is that a small number means a small problem. That one is worse, and the reason is the shape of the curve rather than its height. A rate that multiplies by twelve in about two years, with the inflection arriving alongside the general availability of AI writing tools, is not describing a stable background of sloppiness. The authors are careful about that association and so am I: the audit establishes a coincidence in timing, not the origin of any individual fabricated reference. Nobody looked at those 4,046 citations and confirmed which model, if any, produced them.

What makes even a small number consequential is where it lands. As Topaz put it, a medical professional or clinical guideline developer has no way of knowing that the evidence they are relying on does not exist. Citations are load-bearing. A guideline inherits its authority from the papers underneath it, and nobody re-derives that chain.

The response so far has been thin. By the audit's February cutoff, more than 98% of the affected papers had received no publisher action.

The failure this method cannot see

Now the part that I think matters more than the headline number, and it follows directly from how CITADEL works.

The test is: does a paper with this title exist in any of four databases? That is an excellent test, and it has exactly one blind spot, which is total. It cannot detect a citation that is perfectly real and attached to a claim the cited paper does not make.

Think about what those two errors look like from the outside. A fabricated reference is a dead link. Any librarian, any reviewer, any automated checker can find it in seconds, and once found it is unambiguous. There is no defending a paper that was never written. A misattributed reference is a live link. It resolves. It goes to a real journal, real authors, a real DOI. It looks like diligence. Finding out that it does not support the sentence it is bolted to requires a human being to read the cited paper and form a judgment about what it actually established, which costs perhaps twenty minutes per reference and cannot be automated by a database lookup.

So the detectable error is the rare, obvious, easily corrected one. The undetectable error is the common one, and it long predates any language model. Citing a review as though it were primary evidence, citing a mouse study for a human claim, citing a paper for its abstract's framing when its results section is more equivocal, citing something because three other papers cited it that way and nobody in the chain went back to look: none of that produces a broken link, and none of it would appear anywhere in this audit's 4,046.

"Citation practices are changing with the generative AI use...engagement with the literature is becoming increasingly more superficial."Mohammad Hosseini, Northwestern University, quoted in STAT, 7 May 2026

That is the claim I would keep. Not that AI is filling journals with invented studies, which the numbers do not support, but that the mode of engagement is thinning. Fabricated references are the visible tip of that, and they are visible precisely because they are the version that fails loudly.

Why this error and not another

It is worth being clear about why a reference is such an attractive thing for a language model to get wrong, without pretending to more certainty about mechanism than I have.

A citation is one of the most formulaic strings in all of technical writing. Author surnames, initials, a plausible title assembled from the vocabulary of the field, a journal that publishes that sort of work, a year, a volume, a page range. Every component is highly patterned, and a system optimized to produce text that fits the pattern will produce something that fits the pattern. Whether the referent exists is a separate question that the pattern does not encode.

Which is also why the failure is so easy for a human to skim past. A fabricated citation does not look wrong. It looks exactly like a citation. That is its whole nature. The only way to tell is to go and check, and checking is the step that gets skipped, which brings this back to where the previous reports on this beat kept landing: the problem is rarely the model's output on its own, it is the absence of a verification step that used to be implicit in the work, and the false comfort of a check that does not actually test the thing you care about.

What I would actually do with this

If you write papers, the operative rule is not "don't use AI." It is that a reference you have not opened is not a reference, it is a guess with punctuation. That standard is old and it applies identically whether the guess came from a model, a colleague's slide deck, or your own memory of a paper you read in 2019. Topaz's own framing of the remedy is instructive here, in that he treats one or two incidental fabricated references as a case for correction and transparency rather than retraction. The failure is usually not fraud. It is an unperformed check.

If you read papers, three questions, in descending order of how often they pay off:

Does the citation support the specific sentence? Not the topic, the sentence. This is the check that finds real problems, and it is the only one that requires you to actually read something. Do it for the two or three claims the paper's conclusion actually rests on, and skip the rest.

Does the citation resolve? Paste the title into PubMed or Crossref. Ten seconds. This is now worth doing in a way it was not three years ago, and a dead reference in a recent paper is a legitimate reason to distrust the rest of it, because it tells you the authors did not open their own sources.

Is the citation doing work it cannot do? Watch for a strong claim resting on a single reference, especially a review or a conference abstract. That is a structural weak point regardless of whether the reference exists.

Honest limits

The Lancet correspondence itself is behind a paywall and I did not read it. Everything above comes from sources I did open: the PubMed record for the letter, which confirms its title, authors, DOI and citation details; the Columbia School of Nursing release from the authors' own institution; the CITADEL project page maintained by the lead author; and the independent coverage in Retraction Watch, STAT and The Scientist. All of those agree on the figure of 4,046 fabricated references across 2,810 papers. Retraction Watch gives 4,406, which I take to be a transposition, since the authors' own institution and project page both give 4,046. I have used 4,046 and flagged the discrepancy rather than quietly picking one. It would be a small irony to get a number wrong in a report about not checking your sources.

Two further limits. The corpus is the PubMed Central Open Access subset, which is biomedical and open access, so the rate should not be read as a rate for all of science or all of publishing. And a system that classifies a reference as fabricated when its title appears in none of four databases will inevitably miscount at the edges in both directions, on obscure or non-English or non-indexed work in one direction, and on fabrications that happen to collide with a real title in the other. The authors report a count from an automated pipeline, not a hand-audited census.

The signal

The useful thing about this audit is not the number. It is the demonstration that the citation graph is now something worth auditing at all, which is new, and that the audit was cheap once someone bothered. As Topaz noted about publishers acting on it, the tools exist and the barrier is institutional rather than technological.

But I would not walk away thinking the danger is invented papers. The danger is that the cost of producing something that looks like scholarship fell faster than the cost of checking it, and a fabricated reference is simply the failure mode that happens to leave a mark. The failures that leave no mark did not go anywhere. They got cheaper too.

Which is the whole argument for opening the source. Not because you expect it to be fake. Because you want to find out whether it says what somebody told you it says.

Sources

  1. Maxim Topaz, Nir Roguin, Pallavi Gupta, Zhihong Zhang and Laura-Maria Peltonen, "Fabricated citations: an audit across 2·5 million biomedical papers," The Lancet, 2026 May 9;407(10541):1779–1781, DOI 10.1016/S0140-6736(26)00603-3, PMID 42107362. (The underlying publication, a Letter. Record retrieved from PubMed, which confirms title, all five authors and their affiliations, journal, date, volume, pages and DOI. The full text is paywalled and was not read; no figure in this report is sourced to the letter's own text. Flagged as a limitation above.)
  2. Columbia University School of Nursing, "Nearly 3,000 peer-reviewed medical papers have fake citations, a Columbia Nursing AI-assisted audit finds," 8 May 2026. (Release from the lead author's own institution. Source of 2.5 million papers, 97.1 million references verified, 4,046 fabricated citations across 2,810 papers, the PubMed Central Open Access corpus and the 1 January 2023 to 18 February 2026 window, the author list, and the Topaz statement that a medical professional or clinical guideline developer has no way of knowing that the evidence they are relying on does not exist.)
  3. Maxim Topaz, "CITADEL: Citation Integrity Audit," project page. (Maintained by the lead author. Source of the CITADEL name and expansion, the statement that every reference in the corpus was verified against four independent databases (PubMed, Crossref, OpenAlex and Google Scholar), the 125.6 million reference count, the quarterly rate rising from about four per 10,000 papers in 2023 to 56.9 in early 2026, and the single paper containing 18 of 30 fabricated references.)
  4. Sneha Khedkar, "One in 277 Biomedical Papers Carry Fake References," The Scientist, 2026. (Independent coverage. Source of the method detail that language models compared reference titles against their associated DOIs or PubMed identifiers, that mismatched titles were then cross-referenced against outside databases, and that a reference was classified as fabricated if its title did not appear in any of these databases. Also the Topaz statement that the tools exist and the barrier is institutional, not technological.)
  5. "One in 277 PubMed-indexed papers in 2026 shows fabricated references, says analysis," Retraction Watch, 7 May 2026. (Independent coverage. Source of the per-year rates of one in 2,828 for 2023, one in 458 for 2025 and one in 277 for the first seven weeks of 2026; the mid-2024 inflection; the finding that more than 98% of affected papers had received no publisher action; and the Topaz view that for papers with one or two fabricated references incidental to the main findings, correction and transparency may be more proportionate than retraction. Note: this article gives the fabricated-reference count as 4,406 rather than 4,046, as discussed under Honest limits.)
  6. Anil Oza, "Fraudulent citations, blamed on AI hallucinations, are becoming more common in research papers," STAT, 7 May 2026. (Independent coverage. Source of the blockquote from Mohammad Hosseini of Northwestern University, verbatim, and of the researchers' own acknowledgment that roughly 4,000 fabricated citations across more than 2 million papers is a relatively low proportion.)
Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports