Report 046 · AI in the Lab
Asking three AIs is not a second opinion
It has become a standard move in research groups: put the question to two or three different models, and if they land in the same place, treat that as corroboration. A May 2026 paper measured whether frontier models actually disagree when you ask them for novel hypotheses. Mostly they don't, and the reason is not that they are right.
By Onur Oncer
Published 2026-07-25
Read 7 min
Watch how AI actually gets used in a lab and you will notice the habit almost immediately. Somebody asks a model what to try next. Then, being careful, they ask a different company's model the same thing. When both come back pointing at the same family of candidates, the tension in the room drops. Two independent systems agreed.
That inference is the thing worth examining, because it is doing real work in real decisions about what gets synthesized next week, and it rests on an assumption nobody checks: that the two models are independent witnesses.
The experiment
A team from IIT Delhi and Friedrich Schiller University Jena put that assumption on a scale. The setup, from a paper posted to arXiv on 9 May 2026, is simple enough to describe in a sentence: take six frontier models from two different companies, hand them the same scientific material, and measure how similar their answers are to each other.
Specifically, three Anthropic models (Claude Haiku 4.5, Sonnet 4.5, Sonnet 4.6) and three OpenAI models (GPT-5 Nano, GPT-5 Mini, GPT-5), run against 50 publications from the 2025 NeurIPS AI4Mat track, with ten independent samples drawn per model. Answers were embedded and compared by cosine similarity.
Two tasks, and the contrast between them is the whole design. Task one was interpretive: given a summary of the experiments, recover the hypothesis the paper was actually testing. There is a right answer, so convergence is expected and healthy. Task two was open-ended: given the full paper, propose novel hypotheses. Here convergence is the problem, because six systems generating genuinely creative research directions should scatter.
They didn't. The authors report that inter-model similarities "remain high despite the desired diversity in task outputs." The open-ended generative task produced agreement comparable to the task where agreement was the point.
The obvious objection is that the measurement is broken rather than the models, that the embedding is simply mushing everything into the same region of space. The authors checked that, verifying the embedding model does distinguish content across different papers, which they characterize as confirming a "true lack of epistemic diversity, not degenerate behavior."
Their conclusion is one sentence, and it is the sentence I would tape to a lab wall:
A research community that queries multiple AI systems for novel hypotheses is, from an epistemic standpoint, effectively sampling from a single model.
What that does and does not establish
Fair is fair, so here are the limits, and they are real. This is a position paper, not a peer-reviewed empirical study, and it says so about itself. The experiment covers 50 papers from a single conference track and six models from two vendors. The results are presented as similarity heatmaps rather than one headline statistic, so the honest summary is "similarity stayed high where it should have dropped," not a specific percentage. Someone should run this bigger and across more providers.
What it does establish is that the independence assumption is now a thing you have to defend rather than something you get for free. And there is a plain mechanism behind it. Models trained on overlapping corpora and then tuned by preference optimization get pushed toward the same consensus, a compression effect the paper attributes to prior work by Kirk and colleagues showing these procedures measurably shrink output diversity toward annotator consensus. Two systems shaped by the same literature and the same optimization pressure are not two witnesses. They are one witness with two accents.
The bigger hole: the literature is a highlight reel
The convergence result is the paper's sharpest measurement, but the argument that lands hardest for anyone who has actually run experiments is a different one. The models are trained on published science, and published science is a curated record of things that worked.
The paper puts the tacit part precisely. Laboratory practice, it notes, "produces understanding that is rarely written down": which synthesis conditions are reliable, "which reagents behave inconsistently across suppliers, which reported protocols require undocumented adjustments." Anyone who has tried to reproduce a method from a paper knows this. The written protocol is the skeleton. The knowledge that makes it work lives in a person, and often in a specific person in a specific building.
Then the failure part, which is worse: "Publication bias removes negative results from the corpus." And beyond the missing failed papers, the missing process, since "the iterative cycle of anomaly identification, tentative re-framing, targeted follow-up experiment is not documented even when it eventually produces a publishable result." The paper reports what was concluded. Almost nothing about how anyone got there survives into the corpus.
Run through the paper's own case study, solid-state battery electrolytes, and the consequence is concrete: "Several top-ranked AI candidates for SSEs are known to be unsynthesizable through unpublished tacit knowledge." The model ranks a compound highly. Somebody in a lab already knows you cannot make it. That knowledge exists, it is simply not in any document the model has read.
The 2016 experiment that proves the point in reverse
Here is what makes the failure-knowledge argument more than a complaint. It has been tested, a decade ago, and the result was unambiguous.
In May 2016, a team led by Paul Raccuglia published in Nature a study built on exactly the data everyone throws away. They went into archived laboratory notebooks and pulled out the "dark" reactions: failed or unsuccessful hydrothermal syntheses. They trained a machine-learning model on those, then used it to predict conditions for crystallizing templated vanadium selenites using organic building blocks that had never been tried.
In the paper's own words, the model "outperformed traditional human strategies, and successfully predicted conditions for new organically templated inorganic product formation with a success rate of 89 per cent." Inverting the model, they add, "reveals new hypotheses regarding the conditions for successful product formation." The failures were not noise. They were the most informative data in the building, and they had been sitting in notebooks because there was nowhere to publish them.
Ten years later we are training vastly larger models on the same filtered literature, and the notebooks are still in the drawer.
The signal
None of this says the models are useless in a lab. The same paper is explicit that they already function as capable co-scientists, and my own experience matches: for reading across fields faster than any human can, for drafting an analysis, for catching the thing you stopped seeing three weeks ago, they are excellent. The argument is narrower. They are not built for the autonomous part, and the specific failure mode is that their limitations are correlated, so stacking more of them does not average the error out.
Practically, three things follow.
Stop counting model agreement as replication. A second model is a second draft, not a second opinion. If you want the reassurance that agreement is supposed to provide, it has to come from somewhere the model has not been: your own instrument, your own negative results, or a colleague who has physically attempted the synthesis.
Treat convergence as a warning about coverage. If six models propose the same family of candidates, the useful question is not whether that family is right. It is what is missing from the region they all avoided, because that space is where the unpublished failures and the untried ideas both live.
Write down what did not work. This is the part every one of us controls. Raccuglia's team demonstrated in 2016 what a decade of discarded negative results was worth, and the same paper here proposes a preregistration repository for AI-generated hypotheses so that what a model predicted can be checked against what actually happened. Both amount to the same discipline: stop letting the record be a highlight reel.
This is the same instinct as asking who independently verified an AI result, or noticing when a model nails a protein's known shape and misses the pocket only an experiment could find, or when millions of predicted materials turn out to be predictions. The pattern underneath all of them is that a model is an instrument. Instruments have blind spots. The one thing you must never do with an instrument is check its reading with a second copy of itself.
Sources
- Bisht H, Kumar V, Jablonka KM, Mausam, Krishnan NMA, "Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery," arXiv:2605.08956v1 [cs.AI], 9 May 2026. DOI: 10.48550/arXiv.2605.08956. (Primary. Opened, both the abstract listing and the full HTML text. Position paper from the Yardi School of AI and the departments of Computer Science & Engineering and Civil & Environmental Engineering at IIT Delhi, with the Laboratory of Organic and Macromolecular Chemistry and CEEC at Friedrich Schiller University Jena and HIPOLE Jena. Source for the four stated challenges; for the "hypothesis hivemind" setup quoted above (Claude Haiku 4.5 / Sonnet 4.5 / Sonnet 4.6 and GPT-5 Nano / Mini / GPT-5, 50 publications from the 2025 NeurIPS AI4Mat track, ten independent samples per model, text-embedding-3-small, cosine similarity); for the two-task design; and for every quoted phrase in this report, including "remain high despite the desired diversity in task outputs," "true lack of epistemic diversity, not degenerate behavior," the blockquoted single-model conclusion, "produces understanding that is rarely written down," "which reagents behave inconsistently across suppliers, which reported protocols require undocumented adjustments," "Publication bias removes negative results from the corpus," "the iterative cycle of anomaly identification, tentative re-framing, targeted follow-up experiment is not documented even when it eventually produces a publishable result," and "Several top-ranked AI candidates for SSEs are known to be unsynthesizable through unpublished tacit knowledge." Also the source for the preregistration-repository recommendation and for the attribution of the diversity-compression finding to Kirk et al. Not peer-reviewed; limitations stated in the body above.)
- Raccuglia P, Elbert KC, Adler PD, Falk C, Wenny MB, Mollo A, Zeller M, Friedler SA, Schrier J, Norquist AJ, "Machine-learning-assisted materials discovery using failed experiments," Nature 533, 73–76, May 2016. DOI: 10.1038/nature17439. PMID: 27147027. (Primary. Abstract retrieved and read in full via the Europe PMC record for the DOI; bibliographic details cross-checked against the Semantic Scholar record. Source for the use of "'dark' reactions — failed or unsuccessful hydrothermal syntheses — collected from archived laboratory notebooks," for the templated vanadium selenite target, and for the verbatim result: the model "outperformed traditional human strategies, and successfully predicted conditions for new organically templated inorganic product formation with a success rate of 89 per cent," with inversion of the model revealing "new hypotheses regarding the conditions for successful product formation.")
- The Signal Report, "Did an AI Really Solve an 80-Year Math Problem?," Report 040. (Companion on independent human verification of a model's output.)
- The Signal Report, "The Pocket the AI Couldn't See," Report 003. (Companion: a model correct on the known structure and blind to the cryptic pocket.)
- The Signal Report, "Did AI Really Discover Millions of Materials?," Report 009. (Companion on predicted-but-unmade candidates.)
Onur Oncer
U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.