Two words get used interchangeably in almost every conversation about the credibility of research, including conversations among researchers. Reproducible. Replicable. They sound like synonyms and they are not, and the April Nature package is the cleanest illustration I have seen of why the distinction is worth holding onto.
Here is the difference in one line each.
Reproducible means: I take your data and your analysis, I run it, and I get your numbers. Replicable means: I run your study again on new subjects, and I find what you found.
These are not two grades of the same test. They are tests of different things. Replication asks whether the world behaves the way your paper says it does. Reproduction asks whether your paper describes what you actually did to your own data. A study can be perfectly reproducible and completely wrong about the world. It can also fail reproduction while being right, if the write-up simply does not match the analysis that produced it.
I am a peer-reviewed researcher in microwave spectroscopy, which is about as far from social psychology as you can get and still be doing science. But the reproduction question is field-agnostic. It has nothing to do with subjects or sample sizes. It is the question of whether the arithmetic comes out twice.
The replication number
Tyner and colleagues attempted replications of 274 claims of positive results drawn from 164 quantitative papers published between 2009 and 2018 in 54 journals. This is the paper the headlines were about, and its abstract reports the result plainly:
"Replications showed statistically significant results in the original pattern for 151 of 274 claims (55.1% (95% confidence interval (CI) 49.2–60.9%)) and for 80.8 of 164 papers (49.3% (95% CI 43.8–54.7%)), weighed for replicating multiple claims per paper."
So: 49.3% of papers replicated. Nature's own news coverage carried the headline "Half of social-science studies fail replication test in years-long project," which is an accurate description of this paper. I want to be clear about that, because the problem I am describing is not a bad headline. It is what happens to that headline once it is standing next to a second number that looks like it.
Before leaving this paper, one detail deserves more attention than it got. These were not underpowered replications throwing darts. The abstract notes they were "high powered on average to detect the original effect size (median of 99.6%)," used original materials where available, and were peer reviewed in advance against a standardized protocol. When a replication attempt with a median 99.6% chance of detecting the original effect fails to detect it, the most economical explanation is that the original effect is not the size the original paper said it was.
Which brings up the finding I would put above the headline. The median effect size, in Pearson's r, was 0.25 in the original studies and 0.10 in the replications, which the authors describe as an 82.4% reduction in shared variance. Read that alongside the binary pass-fail rate and the picture changes. This is not a literature that splits cleanly into true findings and false ones. It is a literature whose effects are, on the whole, substantially smaller than published. Some of them shrink past the significance threshold and get counted as failures. Others shrink and stay significant and get counted as successes, while still being much weaker than the paper claimed.
I have written before about why an error bar tells you more than a p-value. This is that argument with a very large sample behind it. The pass-fail verdict is a threshold. The effect size is the measurement. Reporting only the verdict throws away most of what the replication actually learned.
The reproduction number, and its denominator
The companion paper, by Miske and colleagues, asked the more basic question. Its abstract:
"Published claims should be reproducible, yielding the same result when the same analysis is applied to the same data. Here we assess reproducibility in a stratified random sample of 600 papers published from 2009 to 2018 in 62 journals spanning the social and behavioural sciences. The authors of 144 (24.0%, 95% confidence interval (CI) = 20.8–27.6%) papers made data available to assess reproducibility and, for 38 others, we obtained source data to reconstruct the dataset. We assessed 143 out of the 182 available datasets and found that 76.6 (53.6%, 95% CI = 45.8–60.7%) papers were rated as precisely reproducible and 105.0 (73.5%, 95% CI = 66.4–80.0%) were rated as at least approximately reproducible (within 15% of the original effects or within 0.05 of original P values) after inverse weighting each of the 551 claims by the number of claims per paper."
53.6% precisely reproducible. Set next to 49.3% replicable, and you can see how the two collapse into a single vague impression that about half of social science does not hold up.
But look at what 53.6% is a percentage of. The study sampled 600 papers. The authors of 144 of them, 24.0%, made data available. Another 38 datasets were reconstructed from source data. That yields 182 available, of which 143 were assessed. The reproducibility rate is computed on those 143.
Three quarters of the sample never got tested. Not tested and passed, not tested and failed. Untestable, because the data were not there.
That is the finding I would have led with, and it barely surfaced. It also changes how you should read the 53.6%. The papers whose authors handed over usable data are, almost by construction, not a random slice of the literature. They are the ones with organized files, surviving archives, reachable and willing authors, and enough confidence to let a stranger re-run the analysis. If you had to bet on which quarter of a literature is most carefully done, you would bet on that one. The 53.6% is therefore a plausible best case, and the true rate across all 600 is unknown and unlikely to be better.
This is the same structural point as the availability of any measurement: a number computed only where measurement was possible describes the measurable subset, not the population. In my own field the analogous error is quoting a detection rate over the samples that produced a clean spectrum, having quietly dropped the ones that did not.
Why a reproduction failure is worse than it sounds
A failed replication has many innocent explanations. The effect may be real but smaller than first estimated. It may depend on context, on population, on the year. Social behaviour is genuinely variable, and a finding that does not travel is not necessarily a finding that was fabricated.
A failed reproduction has fewer places to hide. Same data. Same stated analysis. Different answer. The world's complexity is not available as an explanation, because the world was not consulted; a file was. What remains is that the description of the analysis does not match the analysis, or that the file provided is not the file used, or that steps were taken which never made it into the methods section.
None of that requires misconduct, and I want to be careful here. Undocumented data cleaning, a filter applied in an earlier session, a software version change, a lost intermediate file: these are ordinary and they are enough. But the practical consequence is the same either way. If the published description of an analysis does not regenerate the published numbers, the methods section is not a specification. And a methods section that is not a specification is the one thing in a paper that everything downstream depends on, including any attempt to replicate it.
That ordering matters. Reproduction is upstream of replication. If you cannot recover a paper's numbers from its own data and its own stated procedure, you do not really know what to run again.
The control group nobody quoted
There is a third paper in the package, and it is the one that turns this from a diagnosis into something actionable. Brodeur and colleagues examined 110 articles from leading economics and political science journals, selected because those journals maintain mandatory data and code sharing requirements.
More than 85% of the published claims were computationally reproducible.
Same broad domain. Overlapping era. Quantitative social science either way. The reproducibility rate is dramatically higher, and the variable that differs is not the intelligence or integrity of the researchers. It is a journal policy about what you must deposit before your paper appears.
Their robustness results are worth noting alongside it, because they are more mixed and I do not want to oversell the good news. When estimates were subjected to alternative reasonable specifications, 72% of statistically significant results stayed significant and in the same direction, with median reproduced effect sizes close to the originally published values. Six independent teams tested pre-specified hypotheses about what predicts robustness. More experienced researchers tended to find lower robustness, and robustness did not track author characteristics or data availability.
So mandatory sharing does not make findings true. It makes them checkable, and checkability is the precondition for everything else. Miske and colleagues found the corresponding trend on the supply side: their paper tracks the percentage of the 62 sampled journals with data sharing, code sharing and reproducibility check requirements from 2003 to 2025, which is the policy variable moving underneath all of this.
Why the counts have decimal points
A small thing that confuses people reading these abstracts: 80.8 of 164 papers, 76.6 papers, 105.0 papers. Papers do not come in tenths.
These are weighted counts. A paper contributing six claims should not get six times the vote of a paper contributing one, so each claim is weighted by the inverse of how many claims its paper contributed, and the paper-level total comes out fractional. Miske and colleagues state it directly: inverse weighting was applied to each of the 551 claims by the number of claims per paper. Tyner and colleagues say their paper-level figure is weighted for replicating multiple claims per paper.
It is a correct choice and worth recognizing on sight, because the alternative quietly hands the loudest result to whichever papers made the most claims.
What I could not confirm
I read the abstracts of all three papers in full and quote them verbatim above. I did not have access to the full texts, so I cannot speak to the coding rules used to grade a reproduction as precise rather than approximate, to how disputes between assessors were settled, or to how the 39 available-but-unassessed datasets in Miske were selected out. Those choices could move the numbers.
The scope is social and behavioural sciences, economics and political science, sampling papers published 2009 to 2018. It is not evidence about chemistry, materials science, or biomedicine, and nothing here licenses extending these rates to those fields. Data-sharing norms differ enormously across disciplines, and given the Brodeur comparison, that difference alone would be expected to matter.
I have also not verified how many total papers the seven-year project touched across all its components; Nature's news coverage describes an initiative spanning 3,900 social-science papers, which I report as their figure rather than one I confirmed against a primary.
The signal
When you next see a claim that half of some field does not hold up, ask which of the two failures is being counted, because the answer changes the meaning completely.
If it is replication, the finding is about the world: the effect is weaker, more conditional, or absent. Expect noise, expect context dependence, and ask for the effect size rather than the verdict. The most useful number in the replication paper is not 49.3%, it is a median r falling from 0.25 to 0.10.
If it is reproduction, the finding is about the paper: the stated analysis does not regenerate the stated numbers. That is a documentation failure, and it is upstream of everything.
And always look for the denominator. In the reproducibility paper, the headline rate was computed on 143 of 600 papers, because that is how many could be checked at all. The most important number in that abstract is 24.0%, the share of authors who made data available. The comparison with journals that require deposit as a condition of publication suggests the fix is not an appeal to individual virtue. It is a rule about what gets deposited, applied by the people who control whether a paper appears.
Sources
- Andrew H. Tyner et al., "Investigating the replicability of the social and behavioural sciences," Nature 652, 143–150, published 1 April 2026. DOI 10.1038/s41586-025-10078-y. (Primary. Abstract opened and read; quoted verbatim above. Source of the 274 claims from 164 quantitative papers published 2009–2018 in 54 journals; the 151 of 274 claims (55.1%, 95% CI 49.2–60.9%) and 80.8 of 164 papers (49.3%, 95% CI 43.8–54.7%) figures; the median 99.6% power figure; the disciplinary range of 42.5–63.1%; and the median Pearson's r of 0.25 for originals versus 0.10 for replications, described as an 82.4% reduction in shared variance. Full text not accessed.)
- Olivia Miske et al., "Investigating the reproducibility of the social and behavioural sciences," Nature, published 1 April 2026. DOI 10.1038/s41586-026-10203-5. (Primary. Abstract opened and read; quoted verbatim above in full. Source of the 600-paper stratified random sample across 62 journals for 2009–2018; the 144 papers (24.0%, 95% CI 20.8–27.6%) with author-provided data plus 38 reconstructed datasets; the 143 of 182 assessed; the 76.6 papers (53.6%, 95% CI 45.8–60.7%) precisely reproducible and 105.0 papers (73.5%, 95% CI 66.4–80.0%) at least approximately reproducible, with the stated tolerance of within 15% of original effects or within 0.05 of original P values; and the inverse weighting of 551 claims by claims per paper. The reference to journal data sharing, code sharing and reproducibility check requirements tracked from 2003 to 2025 is from the paper's Figure 6 caption. Full text not accessed.)
- Abel Brodeur et al., "Reproducibility and robustness of economics and political science research," Nature, published 1 April 2026. DOI 10.1038/s41586-026-10251-x. (Primary. Abstract opened and read. Source of the 110 articles from leading economics and political science journals with mandatory data and code sharing policies; the finding that more than 85% of published claims were computationally reproducible; the robustness figure of 72% of statistically significant estimates remaining significant and in the same direction with median reproduced effect sizes close to originally published values; and the six independent teams testing pre-specified hypotheses, with more experienced researchers identifying lower robustness and robustness not tracking author characteristics or data availability. Full text not accessed.)
- "Half of social-science studies fail replication test in years-long project," Nature news, 1 April 2026. (Secondary, cited for its verbatim headline and its subheading, "Results from massive, 'eagerly awaited' initiative reinforce concerns about the credibility of science — but raise hope for solutions," and for the description of a seven-year project spanning 3,900 social-science papers. Cited as an accurate description of the Tyner replication paper, not as an error.)
Onur Oncer
U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.