← The Signal Report Work with me

Report 092 · Lab Science

The $38 trillion number that was retracted

It was the second most media-covered climate paper of its year, it was cited by central banks, and its authors withdrew it themselves in December 2025. The interesting part is not that a number changed. It is that the mistake underneath it is the same one that ruins undergraduate biology experiments, and it is invisible unless you go looking for it.

Let me put the boundary of this piece up front, because the subject attracts people looking for ammunition and I do not have any to hand out.

This is not a post about whether climate change is expensive. I have no expertise in climate economics and I am not offering an opinion on the underlying question. It is a post about how a specific number acquired a specific confidence interval, why that interval turned out to be too narrow, and what the failure teaches about reading any quantitative result. The statistical issue at the center of it has nothing to do with climate. I have run into the identical problem in a spectroscopy lab, and I have written about it before in the context of cell counts.

The reason it is worth your time is that this particular error is close to undetectable from the outside. You cannot catch it by checking whether the authors were careful, because they were. You cannot catch it by reading the abstract, or the methods, or by noticing a small sample size, because the sample was enormous. That is exactly what makes it a good teaching case.

What the paper claimed

"The economic commitment of climate change," by Maximilian Kotz, Anders Levermann and Leonie Wenz, appeared in Nature on 17 April 2024. It used empirical findings from "more than 1,600 regions worldwide over the past 40 years" to project sub-national economic damages from temperature and precipitation.

Its central result, from the abstract:

we find that the world economy is committed to an income reduction of 19% within the next 26 years independent of future emission choices (relative to a baseline without climate impacts, likely range of 11–29% accounting for physical climate and empirical uncertainty).

The paper put that in currency as "global annual damages in 2049 of 38 trillion in 2005 international dollars (likely range of 19–59 trillion)," and stated that committed damages exceeded the mitigation costs of holding warming to 2 °C "by a factor of approximately six."

Two features of that result made it travel. The word "committed," meaning locked in regardless of emissions choices, and the fact that the horizon was 2049 rather than 2100. It was a large number, it was soon, and it was presented as already decided. According to Retraction Watch, the paper was accessed more than 300,000 times and cited 168 times in Web of Science.

What the retraction says

On 3 December 2025 the authors retracted it. The notice is short, and I am going to quote its operative passage in full rather than summarize, because summaries of this one have been unreliable in both directions.

The authors have retracted this paper for the following reasons: post-publication, the results were found to be sensitive to the removal of one country, Uzbekistan, where inaccuracies were noted in the underlying economic data for the period 1995–1999. Furthermore, spatial auto-correlation was argued to be relevant for the uncertainty ranges. The authors corrected the data from Uzbekistan for 1995–1999 and controlled for data source transitions and higher-order trends as present in the Uzbekistan data. They also accounted for spatial auto-correlation. These changes led to discrepancies in the estimates for climate damages by mid-century, with an increased uncertainty range (from 11–29% to 6–31%) and a lower probability of damages diverging across emission scenarios by 2050 (from 99% to 90%). The authors acknowledge that these changes are too substantial for a correction, leading to the retraction of the paper.

Note what actually moved. The likely range went from 11–29% to 6–31%, which is a widening at both ends, not a collapse. The probability that damages diverge between emission scenarios by 2050 went from 99% to 90%, which is a fall from near-certainty to strong. Retraction Watch reports the revised central estimate as 17% rather than 19%, a figure I could not confirm in a primary source and which I flag as secondary.

So the honest one-line summary is that the central estimate barely moved and the confidence around it moved a great deal. Coverage that reported this as a debunking got it wrong. Coverage that reported it as a minor technical adjustment also got it wrong, because a range that spans 6% to 31% supports very different decisions than one that spans 11% to 29%, and 99% certainty is a different object from 90%.

The notice closes with something I want to give the authors credit for, since this genre of story usually involves people digging in:

The authors appreciate the corrective role of the global scientific community and thank Thomas Bearpark, Dylan Hogan, Solomon Hsiang and Christof Schötz for bringing these issues to their attention. All authors agree to this retraction.

They also posted a corrected version openly, and stated an intent to resubmit for peer review. That is the system working, slowly and in public, which is the only way it ever works.

The part worth learning: more data is not more information

Two separate problems were raised. The Uzbekistan data anomaly is the one that got the headlines, and it is real, but it is also the ordinary kind of problem: bad numbers in a dataset, disproportionate influence on a result, fix the numbers. Bearpark, Hogan and Hsiang published that critique in Nature in August 2025. Their paper sits behind a paywall and I could not open it, so I am not going to characterize its specific findings beyond what the retraction notice itself states.

The second problem is the interesting one, and Christof Schötz posted his version openly on arXiv, so I read all twenty pages of it.

Here is the setup. Most studies of this kind regress economic growth on climate variables using country-year data. The Kotz paper instead subdivided countries into sub-national regions, roughly 20 per country on average, which multiplied the number of observations by about an order of magnitude compared with a country-level panel.

That sounds like an unambiguous improvement, and it is the move the paper was praised for. Schötz's point is that it is only an improvement if the new observations are carrying new information, and he goes and measures whether they are. His framing of the two extremes is the clearest statement of the idea I have read anywhere:

If all subnational regions were fully independent, each would provide unique information, and the effective information would grow proportionally with the number of regions. However, if subnational regions of the same country were perfect copies of one another, no new information would be gained despite the larger dataset.

Then he computes the correlations between the model's residuals. The results are stark. Between different regions of the same country, the average Pearson correlation is 0.65. Within individual large countries it runs higher: 0.79 for China, 0.79 for Brazil, 0.76 for the United States, 0.74 for Russia, and 0.45 for Canada. Regions in different countries show essentially nothing, an average of −0.01, though within the EU it climbs to around 0.30.

Distance does not rescue it. For regions in the same country more than 1,000 km apart, the average correlation is 0.66, statistically indistinguishable from the 0.65 for regions closer than 1,000 km. This is not a geography effect that fades with kilometers. Regions of one country share an economy, a currency, a government and a statistical agency, and their growth rates move together whether they are neighbors or on opposite coasts.

Now the sting. The paper did use clustered standard errors, which is the standard tool for exactly this problem. It clustered by region, which accounts for correlation across years within a region and assumes different regions are independent. And Schötz measures the temporal correlation the paper corrected for: an average of −0.03 across arbitrary years, and 0.03 for consecutive years.

Read those two numbers together. The correlation that was corrected for is about 0.03. The correlation that was not corrected for is about 0.65. The analysis carefully controlled for the negligible dependence and left the large one in.

That is not carelessness, and I want to be clear about that. It is a defensible-looking choice that turns out to be wrong when you go and measure, and almost nobody goes and measures. Schötz notes that spatial correlation in this literature is a known issue with published solutions dating back years, which makes this an oversight rather than an unknown.

Why this is the same error as counting cells in one dish

I wrote a report a few weeks ago on what error bars actually mean, and the worst failure I described there was treating 300 cells counted from a single dish as n = 300. The cells in one dish share a medium, an incubator, a passage number and a bad afternoon. They are not 300 independent observations of biology. They are something closer to one observation, measured 300 times, and calling it n = 300 manufactures confidence that was never in the experiment.

Statisticians call this pseudoreplication. It is the same structure here exactly. Twenty regions in one country are not twenty independent draws from the global relationship between climate and growth. They are closer to one draw, measured twenty times. The standard error shrinks with the square root of your sample size, so inflating the count inflates your certainty, and the inflation is silent. Nothing in the output announces it.

My own field has the same trap wearing different clothes. Take a thousand spectra of one sample on one instrument in one session and you have superb precision on that session. You have one measurement of the substance. The repeat that would actually tell you something is a different sample, a different day, ideally a different instrument, and that repeat is expensive, so the temptation is always to let the cheap replicates stand in for the expensive ones.

This is why I think a climate economics retraction belongs in a lab science column. The domain is irrelevant. The question "are my observations actually independent" is the same question everywhere, it is nearly always answered by assumption rather than by measurement, and it is the single fastest way to be confidently wrong at scale.

Where I stop, and where Schötz goes further

I need to draw a line here carefully, because the two parties do not agree and it would be easy to blur them.

Schötz's own conclusion is considerably stronger than what the authors accepted. He states that correcting the clustering renders the results "statistically insignificant when properly corrected," that with a corrected bootstrap "the first year of discernible damages is shifted from 2049 to beyond the 2100 time horizon," that annual mean temperature "has no significant coefficients," and that corrected model-selection criteria point toward "the trivial model without climate variables being preferred." He also reports that the regression's explanatory power barely moves when the climate variables are added at all: an R² of 0.255 with no climate predictors, and 0.291 with every climate predictor at ten lags each.

The authors did not adopt that conclusion. Their revision widened the interval to 6–31% and kept a 90% probability of scenario divergence, which is a long way from "no detectable effect." Both of these things are true at once: an independent analyst found the result did not survive his correction, and the original authors, correcting the same issue their own way, found that it did in weakened form. That disagreement is live, the revised paper has not been through peer review, and I am not qualified to adjudicate it. Anyone who tells you this retraction settled the underlying science is selling something, in one direction or the other.

What I could not verify

The Bearpark, Hogan and Hsiang commentary in Nature is paywalled and I did not read it. Everything I say about the Uzbekistan issue comes from the retraction notice, which I read in full, and from Retraction Watch. I opened the authors' public replication repository, which contains code and no results. I am not quoting any figure for how much removing Uzbekistan changed the projections, because several such figures are circulating and I could not trace one to a source I had opened.

The 17% revised central estimate is from Retraction Watch, not from a primary document. The corrected version the authors posted is an author correction dated 6 August 2025, openly licensed, and explicitly not peer reviewed. Schötz's arXiv preprint is dated 14 August 2025; a version was published in Nature as a Matters Arising, and my quotations are from the arXiv version I read, not the published one, which may differ.

I have not attempted to reproduce anyone's analysis. I read documents.

The signal

Three things to carry out of this.

First, when a study's strength is that it has a lot of data, ask where the data came from rather than how much there is. Splitting existing units into smaller units multiplies your row count without adding a single new fact about the world. The relevant quantity is never the number of observations. It is the number of independent ones, and that is an empirical question about your residuals, not a property of your spreadsheet.

Second, notice which dependence got corrected. This paper did use clustered standard errors, which is why it looked rigorous, and the cluster definition quietly decided the answer. When you see a robustness measure, find out what it was made robust to. The correction that is present is not evidence about the correction that is absent.

Third, the useful reading of this episode is neither triumphant nor dismissive. Four outside researchers found real problems, the authors agreed, retracted their own highly cited paper, published a corrected version openly, and are resubmitting it for review. The number was too confident and it got fixed in public. If your response to a self-initiated retraction is that science cannot be trusted, you have the sign backwards. This is the part that works.

Sources

  1. Maximilian Kotz, Anders Levermann and Leonie Wenz, "The economic commitment of climate change" (RETRACTED), Nature 628, 551–557, published online 17 April 2024. DOI 10.1038/s41586-024-07219-0. (Primary, opened and read via PubMed Central. Source of the abstract quotation, the 19% and 11–29% figures, the 1,600 regions and 40 years, the $38 trillion 2049 damages with the 19–59 trillion range, and the approximately sixfold mitigation-cost comparison.)
  2. "Retraction Note: The economic commitment of climate change," Nature 648(8094), 764, 3 December 2025. DOI 10.1038/s41586-025-09726-0. (Primary, opened and read in full via PubMed Central; the two long block quotations are verbatim. Source of the Uzbekistan 1995–1999 statement, the spatial auto-correlation statement, the 11–29% to 6–31% and 99% to 90% changes, the too-substantial-for-a-correction language, the acknowledgement of Bearpark, Hogan, Hsiang and Schötz, and the all-authors-agree line.)
  3. Christof Schötz, "Matters Arising: Spatial correlation in economic analysis of climate change," arXiv:2508.10575 [stat.AP], 14 August 2025. Technical University of Munich and Potsdam Institute for Climate Impact Research. Published version: Nature 644, E27–E30, 2025, DOI 10.1038/s41586-025-09206-5. (Primary, 20-page PDF downloaded and read in full. Source of the two-extremes quotation, the correlation table figures including 0.65 same-country, 0.66 above 1,000 km, 0.79 China and Brazil, 0.76 USA, 0.74 Russia, 0.45 Canada, −0.01 different countries and 0.30 EU28, the −0.03 and 0.03 temporal correlations, the roughly 20 regions per country, the R² of 0.255 and 0.291, the clustered-standard-error discussion, the statistically-insignificant conclusion, the 2049 to beyond-2100 bootstrap result, and the model-selection findings. Quotations are from the arXiv version, which may differ from the published Nature version.)
  4. Thomas Bearpark, Dylan Hogan and Solomon Hsiang, "Data anomalies and the economic commitment of climate change," Nature 644, E7–E11, 2025. DOI 10.1038/s41586-025-09320-4. (NOT OPENED. Paywalled; listed for completeness because the retraction notice credits it. No claim in this report is drawn from it. The authors' replication repository at github.com/Global-Policy-Lab/ma_klw was opened and contains replication code without results.)
  5. Ellie Kincaid, "Authors retract Nature paper projecting high costs of climate change," Retraction Watch, 3 December 2025. (Secondary, opened and read. Sole source of the 300,000 accesses and 168 Web of Science citations, the 17% revised central estimate, the August 6 and following-week publication dates of the two commentaries, and the authors' "We broadly agree with the issues raised" statement. Flagged in-text as secondary.)
  6. Maximilian Kotz, Anders Levermann and Leonie Wenz, "Author correction of 'The economic commitment of climate change'," Zenodo, version 1, 6 August 2025. DOI 10.5281/zenodo.15984134. CC BY 4.0. (Record page opened. Cited to establish that the corrected version exists, is openly available and is not peer reviewed. Its numerical contents are not quoted here.)
Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports