An autonomous lab is a loop. Predict what might exist, make it, measure what you got, decide what to try next, repeat without sleeping. When people write about these systems they write about the robot arms, because arms are photogenic. The interesting engineering is in the third step, and the third step is the one I do for a living.
I am a spectroscopist. My job is turning a pattern of peaks into a claim about what a substance is. So when an automated system tells me it identified a compound it had never made before, my first question is not whether the robot worked. It is: how does it know what came out of the furnace?
What the A-Lab reported
In November 2023, a team from UC Berkeley and Lawrence Berkeley National Laboratory, with collaborators at Google DeepMind, published the A-Lab in Nature. It was a genuinely impressive build: a robotic solid-state synthesis line that took candidate compounds from computational screening, proposed recipes using language models trained on the published literature, ran the syntheses, analyzed the products by X-ray diffraction, and used the results to pick its next move.
The paper was titled "An autonomous laboratory for the accelerated synthesis of novel materials." The abstract said this, and I am quoting it as it was published:
"Over 17 days of continuous operation, the A-Lab realized 41 novel compounds from a set of 58 targets."
It closed by saying the success rate "demonstrates the effectiveness of artificial-intelligence-driven platforms for autonomous materials discovery." The coverage wrote itself. A robot chemist had discovered dozens of new materials in under three weeks.
The objection, and why it was a measurement objection
Within about five weeks, chemists pushed back. In January 2024, Robert Palgrave at University College London and Leslie Schoop at Princeton, with Josh Leeman and colleagues, posted an analysis on ChemRxiv arguing that the compounds were not new, and that the identifications were unreliable. As reported by Chemistry World, their conclusion was that no new materials had been discovered in the work.
The technical core of the complaint is the part worth your attention, because it is not a complaint about robots. It is a complaint about pattern fitting.
Powder X-ray diffraction does not show you a compound. It gives you a curve: intensity against angle, a series of peaks whose positions and relative heights encode the repeating geometry of the crystal. To turn that into an identification you use Rietveld refinement. You propose a structural model, compute the diffraction pattern that model would produce, and adjust parameters until the computed curve sits on top of the observed one. When the fit is good, you have shown that your model is consistent with the data.
That is not the same as showing it is the only model consistent with the data, and the gap between those two statements is where this entire story lives. I have made the same argument about what a Raman spectrum can and cannot prove. A fit quality is a statement about agreement, not about uniqueness. A human expert knows to go looking for the other structures that would fit just as well. An automated pipeline fits what it was told to fit.
Palgrave's assessment of the refinements, quoted in Chemistry World, was blunt: it "was very bad, very beginner, completely novice human level." He described the novelty problem as a specific and systematic one: "Two thirds of their entire predictions were just ordered versions of compounds that were already known to be disordered."
That sentence deserves unpacking, because it is the mechanism by which a new compound gets manufactured out of nothing. Many real materials are compositionally disordered: two different atoms share the same crystallographic site, mixed at random. A computational screen that only considers neatly ordered arrangements will generate an ordered version of that same material and, finding no exact match in its reference database, label it new. It is not new. It is a known substance described with a tidier model than reality uses, and its calculated diffraction pattern can look enough like the real one that a fitting routine accepts it without complaint.
What the correction actually says
On January 19, 2026, Nature published an Author Correction from the original team. It is open access, it is short, and I recommend reading it yourself. Three things changed.
First, the novelty claim was reframed. In the authors' words: "We acknowledge that the original claims of material novelty were subject to misinterpretation, their intention was to indicate that the materials were new to the prediction platform, not necessarily new to science."
Second, the identifications were redone by hand. "We have manually re-analyzed the diffraction patterns and have confirmed that the prediction platform came to the correct conclusion in 36 of its 40 reported successes, with 4 compounds being inconclusive." The correction notes that this re-analysis was itself peer reviewed after publication, and Nature thanks both the correspondents who raised the issues and the reviewers who assessed the re-analyzed data.
Third, and this is the one that made me sit up: "We have also removed one compound from the discussion (Zn2Cr3FeO8) that was mistakenly included in the training data."
Read the diff
The reason this case is worth a whole report is that the correction is legible. You can put the old abstract next to the new one and see, word by word, what a scientific claim looks like when it is trimmed back to what the data support.
The title went from "the accelerated synthesis of novel materials" to "the accelerated synthesis of inorganic materials."
The headline sentence went from "the A-Lab realized 41 novel compounds from a set of 58 targets" to "the A-Lab realized 36 compounds from a set of 57 targets." The word "novel" is simply gone.
And the closing sentence went from "artificial-intelligence-driven platforms for autonomous materials discovery" to "for autonomous materials synthesis."
The arithmetic is traceable. Drop the one compound that had leaked into the training data and you go from 58 targets to 57, and from 41 successes to 40. Re-analyze those 40 by hand, find 4 that diffraction alone cannot settle, and you are at 36. Nothing here is hidden.
But look at the last edit again, because it is doing more work than the numbers. Discovery became synthesis. That is the accurate description of what an autonomous lab does, and it is a smaller claim than the one the world heard.
"Novel" was never a chemistry claim
Here is the part I would most like people to carry away, because it generalizes far past this one paper.
When a computational pipeline calls a compound new, it is not making a statement about chemistry. It is making a statement about a database. It means: I checked my reference set and did not find this. Whether that is a discovery depends entirely on how complete the reference set is, how the entries are represented, and whether the thing you are searching for would be recognizable if it were already there. Change any of those and the same compound flips between new and known without a single atom moving.
I made this argument in a different context when an AI predicted 2.2 million crystals and working chemists found little novelty or usefulness in a sample of them. That report was about prediction: compounds nobody had made. This one is the sequel that matters more, because here the compounds were made, in a real furnace, and the question of what they were still came down to how a pattern got interpreted. Automating the making did not automate the knowing.
The training-data leak is the other lesson, and it is an old one wearing a lab coat. A model that has seen the answer will reproduce the answer, and if you count that as a success you have measured your own bookkeeping. Nobody did this on purpose. That is precisely why it is worth naming: in a loop running unattended for 17 days, the contamination check is not a step anyone is standing at.
What still deserves credit
I want to be fair to the work, because a correction is not a retraction and the coverage of these disputes tends to flatten everything into fraud or vindication.
Thirty-six confirmed solid-state syntheses in 17 days of unattended operation is a real result. The recipe proposal, the robotics, the closed-loop retry logic all worked at the job they were built for. The authors also did the thing you want people to do: they re-analyzed by hand, they submitted the re-analysis to review, they corrected the record in place and published an annotated version of the original showing every textual change. That is more transparency than most corrections carry, and the machinery of post-publication review functioned, if slowly. Twenty-six months slowly.
Two honest limits on what I have told you. The re-analysis that produced the 36 figure was performed by the original authors, though Nature states it was peer reviewed post-publication. And the critics' strongest claim, that no new materials were discovered at all, is their conclusion, not something the correction concedes in those words. I could not open the ChemRxiv preprint directly, so I have relied on Chemistry World's reporting of it and quoted only what that piece quotes.
The signal
An autonomous lab can only be as autonomous as its weakest measurement. The A-Lab could synthesize faster than any group of humans could characterize, which sounds like a triumph until you notice what it implies: the identification step becomes the throughput limit, so the pressure is to automate that too, and once you do, nobody is checking anything. The loop closes on itself. Speed was never the scarce resource in materials science. Certainty was.
This is the same shape as the two Erdős claims seven months apart, where the variable that decided which one survived was not model capability but who checked the work. And it is why I keep coming back to the same unglamorous position: a reading is not a result until someone competent has asked what else could have produced it. That question is cheap to ask and expensive to skip, and no amount of throughput answers it for you.
The most useful sentence in this whole affair is one the authors wrote themselves: the materials were new to the prediction platform, not necessarily new to science. Every AI discovery headline you read this year would be improved by having that clause appended to it.
Sources
- N. J. Szymanski, B. Rendy, Y. Fei, R. E. Kumar, T. He, D. Milsted, M. J. McDermott, M. Gallant, E. D. Cubuk, A. Merchant, H. Kim, A. Jain, C. J. Bartel, K. Persson, Y. Zeng and G. Ceder, "Author Correction: An autonomous laboratory for the accelerated synthesis of inorganic materials," Nature 650(8100):E1, published online January 19, 2026, DOI 10.1038/s41586-025-09992-y. (Primary, open access, read in full. Source of every verbatim quotation about the correction: the misinterpretation of the novelty claim, the "36 of its 40 reported successes, with 4 compounds being inconclusive" re-analysis, the removal of Zn2Cr3FeO8 as "mistakenly included in the training data," and the statement that the re-analysis was peer-reviewed post-publication.)
- Same authors, "An autonomous laboratory for the accelerated synthesis of novel materials," Nature 624(7990):86–91, published online November 29, 2023, DOI 10.1038/s41586-023-06734-w, PMID 38030721. (The original record, retaining the pre-correction title and abstract. Source of the verbatim original wording: "realized 41 novel compounds from a set of 58 targets" and "autonomous materials discovery." Cross-checked against the U.S. Department of Energy's OSTI record for the same article, which carries the identical original abstract.)
- The corrected article as served by PubMed Central (PMC10700133). (Source of the post-correction wording used for the side-by-side comparison: the title "of inorganic materials," the abstract sentence "realized 36 compounds from a set of 57 targets," and the closing phrase "autonomous materials synthesis." The record also carries the correction notice pointing to Nature 2026 Jan 19;650(8100):E1.)
- "New analysis raises doubts over autonomous lab's materials discoveries," Chemistry World, January 16, 2024. (Source for the existence and content of the ChemRxiv critique by Josh Leeman, Leslie Schoop, Robert Palgrave and colleagues, and for both Palgrave quotations used above. The ChemRxiv preprint itself returned HTTP 403 to every fetch attempt during verification, so nothing is cited from it beyond what this opened article reports.)
Onur Oncer
U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.