I make my living on measurement. My own research is microwave spectroscopy, which is a field where you spend most of your time deciding whether the feature in front of you is a property of the sample or a property of the instrument. That habit is why this story caught me, and it is the only expertise I am claiming here. I am not a scientometrician and I have never computed a CD index.
But the failure being argued about in Nature this month is one I recognize immediately, because it is the same failure I have written about with XPS charge referencing and with the 260/280 ratio. A number is computed from a measurement. The measurement can fail in a specific way. When it fails that way, the number does not go blank or throw an error. It returns a clean, plausible, extreme value, and that value goes into the analysis looking exactly like data.
What the original paper claimed
Michael Park, Erin Leahey and Russell Funk published "Papers and patents are becoming less disruptive over time" in Nature on 4 January 2023 (volume 613, pages 138 to 144). They applied a measure called the CD index to roughly 45 million papers and 3.9 million patents, and reported a decline in the share of work that is disruptive rather than consolidating.
The CD index, roughly, asks what happens to a paper's own references after that paper appears. If later work cites the paper but stops citing what the paper built on, the paper displaced its predecessors and scores toward +1, disruptive. If later work cites the paper alongside its predecessors, it consolidated an existing line and scores toward -1. It is a genuinely clever idea, and its appeal is that it measures something about the structure of the citation network rather than counting citations.
The reception was enormous. VUB's account, from the team that wrote the critique, records over 300 press stories and more than a thousand citations in the scientific literature, and coverage of the exchange notes at least sixteen policy documents referencing the original. The finding became a premise. If you have read an essay in the last three years arguing that science is stagnating, that ideas are getting harder to find, or that funding structures have gone conservative, there is a good chance this paper was sitting underneath it.
The failure mode
Now the critique. Vincent Holst, Andres Algaba, Floriano Tori, Sylvia Wenmackers and Vincent Ginis, at the Data Analytics Laboratory of the Vrije Universiteit Brussel and the Centre for Logic and Philosophy of Science at KU Leuven, published "Dataset artefacts can partially drive the measured decline in disruption" in Nature on 12 August 2026, as a Matters Arising (volume 656, pages E7 to E13).
Their preprint states the mechanism compactly. This is from the abstract of the February 2024 arXiv version:
Due to a factual plotting mistake, database entries with zero references were omitted in the CD index distributions, hiding a large number of outliers with a maximum CD index of one, while keeping them in the analysis. Our reanalysis shows that the reported decline in disruptiveness can be attributed to a relative decline of these database entries with zero references.
Take it in two pieces, because they are separate problems and only one of them is about a plot.
The first piece is the metric's behaviour on degenerate input. If a record lists no references, then there are no predecessors for later work to cite, so by construction nothing can be found to consolidate, and the index returns its maximum: 1, perfectly disruptive. That is not a bug in anyone's code. It is what the formula says. But it means the value 1 is doing double duty: it marks the genuinely field-founding paper, and it marks the record whose reference list is simply absent from the database.
The second piece is that these records are not evenly spread across time. Older material has worse metadata. Reference lists for a 1955 paper are far more likely to be missing from a citation database than for a 2015 paper, because the digitization and parsing that produced those records got better over the decades. So the population of artefact-driven 1.0 scores shrinks over time. A shrinking population of maximum scores, plotted against year, looks exactly like a decline in disruptiveness. The critics' claim is that what was being measured was substantially the improvement of the database.
Their strongest supporting point is the one that makes this an artefact rather than a judgement call: they report that the source documents for these zero-reference entries predominantly do make references. The references exist in the papers. They are missing from the records.
The preprint adds two consequences worth noting. The robustness checks in the original did not catch it, because the regression adjustment cannot control for the hidden outliers, which sit at a discontinuity in the index rather than in a tail. And the Monte-Carlo simulations, properly evaluated, reproduce the observed decline from random citation behaviour once the outliers are preserved. A control that reproduces your effect is a control telling you something.
The headline number, from the published version as VUB and the authors describe it: in the largest dataset, the artefacts account for 93 percent of the reported decline between 1945 and 2010.
Read the two titles next to each other
There is a detail here that I think is more instructive than the statistics, and it is free to observe.
The preprint, posted to arXiv on 7 February 2024, is titled "Dataset Artefacts are the Hidden Drivers of the Declining Disruptiveness in Science." The version Nature published on 12 August 2026 is titled "Dataset artefacts can partially drive the measured decline in disruption."
Are the hidden drivers, to can partially drive. Declining disruptiveness in science, to the measured decline in disruption. Every claim in that title got narrower, and the second version is careful to say that what declined is a measurement rather than a property of science.
I flag this because it is a live problem for anyone citing preprints, and I have been caught by it here before. The version of a paper you can read for free is often not the version that survived review, and the difference is usually in exactly the direction that matters: the free one claims more. Cite the version you actually read, say which one it was, and check the title.
The reply, and what is not settled
Park, Leahey and Funk published a reply in the same issue (Nature volume 656, pages E14 to E21, 12 August 2026). Nature's Matters Arising format puts both in front of the reader at once, which is the correct way to handle this and is worth appreciating.
I could not obtain the full text of either the Matters Arising or the Reply. Both are behind the paywall and I am not going to characterize an argument I have not read. What I can report, from coverage that I did read, is the shape of their response: they argue the reported decline remains statistically significant under alternative specifications, and they raise concerns about the composition of the reanalysis dataset. If you want to know who is right, you need both papers, and I would rather tell you that than paraphrase a rebuttal from a summary.
What I will report with confidence is what the critics themselves say their result does not show, because this is the part the coverage mangled. One outlet headlined it as science not becoming less disruptive after all. That is not the claim. Ginis, quoted in the EurekAlert release from his own institution:
Based on the used measure, we simply cannot make conclusive statements about whether science is getting more or less disruptive.
And in the VUB release, more bluntly: "That does not mean we can now conclude that science is actually becoming more groundbreaking."
That is the honest position and it is less satisfying than either headline. A contaminated instrument does not tell you the opposite of what it said. It tells you that you do not have the reading yet. Removing an artefact from a measurement recovers uncertainty, it does not flip the sign. Anyone now citing this exchange as proof that science is fine is making the identical error as citing the 2023 paper as proof that it is not, one reversal downstream.
Thirty-two months
The last thread is the one with the clearest institutional lesson, and the dates are on the PubMed record rather than in anyone's press release.
The Matters Arising was received on 29 October 2023, roughly ten months after the original appeared. It was accepted on 9 June 2026 and published on 12 August 2026. That is about thirty-two months from submission to acceptance, during which the paper being questioned continued to be cited, covered and written into policy. Coverage of the exchange puts the additional citations accrued in that window at over a thousand.
Holst describes how he found it, quoted in the phys.org account:
Shortly after its publication, I attempted to reproduce the study using open data. I discovered a large peak of papers receiving the maximum disruption score. This was strange because the histograms published in Nature showed no such peak.
That is a good scientist doing an ordinary thing: recomputing a published result and noticing that his distribution did not match the printed one. The discrepancy was visible within months of publication. Printing it took years.
Ginis draws the comparison himself, and it lands:
Whereas the software bug was fixed by the open-source community in months, it took years for the scientific correction to be published.
I do not think the useful reaction is to be angry at Nature. Matters Arising is a real process with real standards, the original authors had to be given a chance to respond, and the journal did publish both. The system worked. It worked at a speed that does not match the speed at which a result propagates, and that mismatch is the structural fact, not a scandal. A finding travels in weeks. A correction to it travels in years. Anything built in between is built on the uncorrected version.
This is the same shape as the $38 trillion climate number, where a heavily cited figure was withdrawn by its own authors long after central banks had used it. In both cases the correction was eventually made, in public, by the normal machinery. In both cases the corrected version will never catch the original.
What to take from this
If you use citation metrics for anything consequential, hiring, funding, evaluation, the operational lesson is narrow and concrete: find out what your metric returns when its input is missing. Not what it returns for a weak paper. What it returns for a record where the field is empty. If the answer is a clean extreme value rather than an error, you have a pipeline that converts missing data into strong evidence, and you will not see it in a histogram if the plotting code drops the degenerate cases.
That generalizes past bibliometrics, which is why I wrote it up. A spectrometer with nothing in the beam does not report "no sample." It reports a spectrum. The discipline of asking what your instrument says when you feed it nothing is the cheapest error-detection there is, and in this case one person doing it with open data found in months what took the literature three more years to print.
What I did not verify
I did not read the full text of either the Matters Arising or the Reply. Both are paywalled. My verified sources for those two are the PubMed and Crossref records, which I retrieved directly through the NCBI and Crossref APIs and which establish the titles, authors, affiliations, journal, volume, page ranges, DOIs, and the received and accepted dates. Everything I state about the mechanism comes from the authors' own arXiv preprint, which I read in full, and I have flagged above that the preprint's claims are stated more strongly than the published title.
The 93 percent figure is not from a source I opened directly. It appears in the EurekAlert release issued by the authors' institution, which I read, and I quote it as their characterization of their published result. I did not see it in the preprint and I have not confirmed which dataset "largest" refers to or how the 1945 to 2010 window was chosen. Treat it as the authors' summary of their own paper rather than as a figure I checked.
My account of what Park, Leahey and Funk argue in reply comes entirely from the phys.org article, which I read. It contains no direct quotations from the Reply and no figures from it. I have kept my description at that level deliberately. If you are evaluating this dispute seriously, read both papers.
I have also not independently assessed whether the CD index is a good measure of disruption in the first place. There is a substantial separate literature arguing it is confounded by citation inflation and by changes in citation practice, which I noted exists but did not read for this report. That question is upstream of this dispute and is not resolved by it.
Finally, this is a live exchange. If either team publishes further analysis, or if a correction or addendum issues on any of the three papers, this report needs a dated update and will get one.
Sources
- Vincent Holst, Andres Algaba, Floriano Tori, Sylvia Wenmackers and Vincent Ginis, "Dataset artefacts can partially drive the measured decline in disruption," Nature 656, E7–E13 (12 August 2026), DOI 10.1038/s41586-026-10787-y, PMID 42587124. (The critique. Paywalled; not read in full. Bibliographic record and the received 29 October 2023 / accepted 9 June 2026 dates retrieved directly from the NCBI E-utilities and Crossref APIs, which is the source for the thirty-two month figure and for the authors' VUB and KU Leuven affiliations.)
- Vincent Holst, Andres Algaba, Floriano Tori, Sylvia Wenmackers and Vincent Ginis, "Dataset Artefacts are the Hidden Drivers of the Declining Disruptiveness in Science," arXiv:2402.14583v1 [cs.DL], submitted 7 February 2024. (Open preprint of the above, read in full. Source of the blockquote on the plotting mistake and zero-reference entries, the point that the source documents predominantly do make references, the regression-adjustment and discontinuity argument, and the Monte-Carlo point. Note the title and hedging differ from the published version, as discussed above; where the two differ, the published version governs.)
- Michael Park, Erin Leahey and Russell J. Funk, "Reply to: Dataset artefacts can partially drive the measured decline in disruption," Nature 656, E14–E21 (12 August 2026), DOI 10.1038/s41586-026-10788-x, PMID 42587120. (The reply. Paywalled; not read. Listed so readers can go to it directly. Bibliographic record verified via NCBI E-utilities and Crossref.)
- Michael Park, Erin Leahey and Russell J. Funk, "Papers and patents are becoming less disruptive over time," Nature 613, 138–144 (4 January 2023), DOI 10.1038/s41586-022-05543-x, PMID 36600070. (The original paper. Bibliographic record and the received 14 February 2022 / accepted 8 November 2022 dates verified via NCBI E-utilities. The 45 million papers and 3.9 million patents figures are as characterized in the critique's abstract.)
- Vrije Universiteit Brussel, "VUB research calls the global debate on scientific innovation into question," press release, August 2026. (Institutional release from the critics' own university, read in full. Source of the 300-plus press stories and 1,000-plus citations figures, and of the Ginis quotation that the result does not license concluding science is more groundbreaking.)
- "Science not becoming less disruptive after all," EurekAlert news release, 12 August 2026. (Read in full. Source of the 93 percent figure quoted above and of the Ginis quotation on what the measure can and cannot conclude. Its headline is the overreach discussed in the body.)
- "Claim that science is becoming less disruptive challenged after years-long publication delay," Phys.org, August 2026. (Coverage, read in full. Source of the Holst quotation on reproducing the study and finding the peak, the Ginis quotation comparing the software fix to the scientific correction, the characterization of the Park, Leahey and Funk reply, and the sixteen policy documents figure.)
Onur Oncer
U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.