There is a step in every "AI discovers new materials" story that nobody photographs. Before the model there is a table, and before the table there is a graduate student at a bench at eleven at night, running reaction number forty-one. The model is downstream of that person. Everything it knows, it knows because somebody measured it.
Which raises a question the field has mostly waved at rather than answered: if two labs run the same experiment, do they get the same number? And if they do not, what happens to a model trained on both?
Four laboratories went and checked. The result appeared in Nature Catalysis on 31 July 2026, from a team spanning SLAC National Accelerator Laboratory, Penn State, Stanford and UC Santa Barbara, and it is filed as an Analysis rather than a research article, which is the right shelf for it. This is a paper about measurement, not about a discovery.
The design, which is the whole point
They ran a round-robin study: the same materials and the same agreed protocol, executed independently at multiple sites, with the differences between sites treated as the object of study rather than as noise to be apologized for.
The chemistry was rhodium on titania, Rh/TiO2, running CO2 hydrogenation. Feed carbon dioxide and hydrogen over the catalyst and you can get carbon monoxide, which here is the product you want, or methane, which here is the side product you do not. How the catalyst splits between those two is its selectivity, and selectivity is exactly the kind of number a discovery model is asked to predict.
The inputs they varied were reaction temperature, rhodium loading and synthesis method. The outputs they measured were conversion, selectivity, and the production rates of CO and CH4. Ordinary catalysis. Nothing exotic, which is what makes the finding land.
The finding
Inside any single laboratory, the relationships behaved. Change the temperature, the outputs moved in a way you could fit. Change the loading, likewise. That is the experience every experimentalist has and every dataset is built on.
Then they pooled the four labs, and the paper reports that those same input-output relationships "became statistically insignificant when interlaboratory variability was included." Identical catalyst batches. Identical testing protocols. The effects did not reverse, or turn noisy in some subtle way that needs a statistician to adjudicate. They stopped clearing significance.
On the cause, the abstract is terse: "Heat management emerged as a key contributor to this variability."
I want to flag a discrepancy here rather than paper over it, because it is instructive in its own right. The institutional press releases from Penn State and from EurekAlert describe the dominant difference as stirring intensity, how hard the reaction mixture was agitated. The paper's own abstract says heat management. These are not in conflict, since how vigorously you stir a reactor governs how heat and reactants move through it, and stirring is one of the levers by which heat management succeeds or fails. But they are different levels of description, and the press version is the more specific claim. I am reporting the paper's language as the paper's, and the stirring detail as the institutions' gloss on it, because I could not read the body of the paper to see which the authors emphasize. More on that limitation below.
The authors' own framing, from the Penn State release:
An AI model is only as good as the data used to create it, which will undoubtedly come from multiple sources.
— Robert Rioux, Friedrich G. Helfferich Professor of Chemical Engineering, Penn State
Stirring is a real variable, and somebody measured that too
If "the stir rate mattered" sounds like an excuse rather than a finding, there is a separate paper worth knowing about, and it happens to sit in this one's own bibliography.
In JACS Au in June 2025, a group at the Zelinsky Institute published work under the plainest possible title: magnetic stirring may cause irreproducible results in chemical reactions. Their observation is that the rotation of a stir bar changes depending on where the vessel sits on the stirrer plate, and at certain positions on the plate it stops entirely. They then ran parallel reactions in vials standing next to each other on that same stirrer, and report that conversions in a catalytic cross-coupling differed significantly between them. For nanoparticle synthesis they saw differences in particle morphology and process rate depending on vessel position.
Same room. Same stirrer. Same operator. Different square inch of the plate.
Their proposed fix is refreshingly cheap: run a control experiment with a single vessel placed in the most appropriate position, usually the center of the stirrer, and report that result alongside the parallel ones. One extra vial, and the positional artifact becomes visible instead of silent.
This is the same shape as the microplate story I covered last week, where position in a 96-well plate was worth up to 35 percent of the signal. Position in an apparatus is a variable whether or not your analysis has a term for it. Wells in an incubator, vials on a stir plate, reactors in four different buildings.
Why this is specifically an AI problem, and not just a chemistry problem
Interlaboratory variation is old news. Metrology has had a formal vocabulary for it for decades: repeatability is the spread you get repeating a measurement under the same conditions, same operator, same instrument, short interval, and reproducibility is the spread you get when those conditions change. Round-robin studies exist precisely to quantify the second one. Any lab that has ever run a proficiency test or calibrated against a reference standard has met this idea.
I spent years doing microwave spectroscopy, and in that world you internalize early that the instrument is part of the result. Two spectrometers do not agree until you make them agree, and the making is most of the work. So the finding here does not surprise me. What is new is the consumer.
Traditional catalysis absorbed interlab scatter through a human filter. A reviewer who has run these reactions reads a paper, thinks "that selectivity looks high for that loading," and asks a question. A model does not do that. A model fits whatever structure is present in the table it is given, and it cannot tell which structure came from chemistry and which came from a building.
That is the failure mode worth naming. If a pooled dataset contains a systematic per-lab offset, a sufficiently flexible model will learn the offset. It will use lab identity, encoded implicitly in whatever correlates with it, as a predictor. On a random train-test split it will look excellent, because a random split puts every lab on both sides of the split, so recognizing the lab is rewarded at test time exactly as much as understanding the chemistry.
That last point is mine, not the paper's, and it is the practical consequence I would most want a group to act on: if your data came from multiple sources, do not evaluate on a random split. Hold out an entire lab, train on the rest, and predict the one you held out. That number is the one that tells you whether you learned chemistry. It will be worse. That is the point of it.
Readers of this beat will recognize the pattern. It is the same one as the model that learned which doctor ordered the test rather than the disease, and a cousin of the models that cannot beat a straight line once you check properly. In every case the model found a shortcut that was really in the data, and the data was really about something other than what the researchers thought they were studying.
What the authors ask for
Their recommendations, as reported by the institutions, are about consistency: reactor design, operating protocols, experimental conditions. The abstract's own conclusion is broader and, I think, more useful, arguing that uncertainty analysis has to enter at the stage where you choose performance metrics and input features for a model, not afterward as an error bar bolted onto a prediction.
That ordering matters. Uncertainty as a post-hoc decoration tells you how confident a model is about a relationship it has already decided exists. Uncertainty as a feature-selection criterion asks a prior question: given how much this quantity moves between labs, is it a thing a model can learn from at all? Some of the features in this study did not survive that question.
Our findings are a reminder to exercise caution about what information we feed into a machine-learning model.
— Selin Bac, postdoctoral researcher at UC Santa Barbara and first author
What I am not claiming, and what I could not read
Plainly: the body of this paper is paywalled and I did not read it. What I retrieved and read was the full abstract, the figure captions, the reference list and the acknowledgements from the publisher's page, plus the Penn State and EurekAlert releases. Every specific claim above traces to one of those. I have no per-lab numbers, no effect sizes, no p-values, and no description of the statistical model, because those live in the part I could not open. If you are going to act on this, buy or borrow the paper.
The underlying data appears to be deposited on Zenodo, which is the right thing for the authors to have done. I could not retrieve it either; the repository refused my requests. So I cannot independently confirm anything from the raw data.
This is one round-robin, on one catalyst system, at four labs that are all well-resourced US research institutions working in coordination. It is a lower bound on interlaboratory variation, not an average. Four labs that agreed on a protocol in advance and shipped each other identical material are far more aligned than four labs whose data you scraped from four separate papers.
And a boundary on the general claim: none of this says pooled multi-lab data is useless, or that machine learning for catalysis does not work. It says the uncertainty in that data is usually unmeasured, and that when this group measured it, it was large enough to erase effects. Those are different statements, and the second one is the one the study supports.
The signal
Three things to carry out of this.
First, an experimental dataset is not a set of facts. It is a set of measurements, each carrying an uncertainty that mostly went unrecorded. Pooling data from many sources does not average that away, it adds a new source of variation that lives between the sources and does not shrink as you add rows.
Second, the scaling instinct fails here in a specific way. More data usually helps a model. More data from more labs, without accounting for which lab, can make a model more confident and less right at the same time, because it is learning a real structure that happens not to be chemistry.
Third, and this is the part I would put on the wall: the oldest tool in measurement science, having a second lab run your experiment, is now a machine-learning tool. Not because machine learning changed, but because the data pipeline got long enough that nobody at the model end can see the bench at the other end. The round-robin is how you look.
Sources
- Selin Bac, Dongjae Shin, Seunghwa Hong, Jake Heinlein, Anastassiya Khan, Greg Barber, Zhihengyu Chen, Michael M. Albrechtsen, Christopher Tassone, Robert M. Rioux, Matteo Cargnello, Simon R. Bare, Kirsten Winther, Phillip Christopher and Adam S. Hoffman, "Quantifying uncertainty in catalyst activity and deactivation during CO2 hydrogenation via round-robin testing for data-driven modelling," Nature Catalysis 9(8):912-923, published online 31 July 2026. DOI 10.1038/s41929-026-01559-y. (Primary source. The article body is paywalled and was not read. The abstract, figure captions, reference list and acknowledgements were retrieved from the publisher's page and are the source of the following: the four-laboratory round-robin design on Rh/TiO2 for CO2 hydrogenation; the inputs (reaction temperature, Rh loading, synthesis method) and outputs (conversion, selectivity, CO and CH4 production rates); the quoted finding that relationships clear intralaboratory "became statistically insignificant when interlaboratory variability was included"; the quoted attribution that "Heat management emerged as a key contributor to this variability"; the conclusion that uncertainty analysis must enter at metric and feature selection; the article's classification as an Analysis; and the author list and pagination. Publication date, volume and pagination confirmed independently against Crossref.)
- Penn State University, "Variations between labs can misinform scientific AI models, team reports," August 2026. (Institutional release, opened and read. Source of the verbatim Robert Rioux and Selin Bac quotes and their titles, of the participating institutions, of the description of the dominant difference as stirring intensity, and of the recommendations on reactor design and protocol consistency. The release text was cross-checked against the EurekAlert version below and matched.)
- EurekAlert (AAAS), "Variations between labs can misinform scientific AI models, team reports," release 1140027, August 2026. (Opened and read as a cross-check on the release text. Source of the funding attribution to the US Department of Energy Office of Science, award FWP 101064, and independent confirmation of the 31 July 2026 publication date and the DOI.)
- Veronika A. Cherepanova, Evgeniy G. Gordeev and Valentine P. Ananikov, "Magnetic Stirring May Cause Irreproducible Results in Chemical Reactions," JACS Au 5:3789-3798, published 11 June 2025. DOI 10.1021/jacsau.5c00412, PMC12381720. Open access. (Full text retrieved via the Europe PMC REST API and read. Source of the finding that stir-bar rotation changes with vessel position on the stirrer plate and ceases entirely at some locations, of the significantly different conversions between cross-coupling reactions in adjacent vials on one stirrer, of the nanoparticle morphology and process-rate differences by position, and of the proposed single-vessel centered control experiment. Found via the Nature Catalysis reference list.)
- Not opened, and flagged for transparency: Bac, S. et al., "Dataset: quantifying uncertainty in catalyst activity and deactivation during CO2 hydrogenation via round-robin testing for data-driven modeling," Zenodo, 2026. DOI 10.5281/zenodo.20370161. (The authors' deposited dataset, listed in the paper's reference list. The repository returned HTTP 403 to my requests, so the raw data was not retrieved and no claim in this report rests on it.)
Onur Oncer
U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.