← The Signal Report Work with me

Report 136 · AI in the Lab

AI predicted the crystal. Not the disorder.

A group at the Fritz Haber Institute, Imperial College London and Bayreuth trained a classifier to answer one narrow question: is this composition the kind that comes out of a furnace with its atoms neatly sorted? Pointed at two of the big computational materials databases, it produced two numbers. The coverage carried one of them, without the sentence the authors attached to it.

A predicted crystal structure is a list. These atoms, at these positions, in this repeating box. The energy calculation that decides whether the structure is stable takes that list literally, because it has to: every site holds exactly one kind of atom, and the pattern repeats forever without a mistake in it.

Real solids are frequently not like that, and the gap is not an edge case. It has a name, and someone has now tried to measure how much of the AI-generated materials landscape falls into it.

What disorder means here

The paper restricts itself to one kind, which is the right move and worth stating precisely. Substitutional disorder is, in the authors' words, "statistical chemical mixing on one crystallographic site, where two or more different species share that site with fractional occupancies."

Two elements that resemble each other, sitting on the same position in the lattice, and which one you get at any given spot is a matter of statistics rather than of pattern. There is no single unit cell you can draw that is the truth. Averaged over the sample, that site is 60% one element and 40% the other, and the material's real properties follow from the mixture and from how it is arranged locally, not from the tidy ordered compound with the same formula.

This matters for the same reason it is easy to skip. A high-throughput screen or a generative model proposes an ordered arrangement, computes its energy, finds it sits on or near the convex hull, and files it as a stable new compound. If the material that actually forms is a disordered solid solution, the entry is not exactly wrong. It is a description of something that does not separately exist, and the properties predicted from it can be badly off.

The instrument they built

The team trained on the Inorganic Crystal Structure Database, the accumulated record of experimentally determined inorganic structures, curated down to 118,684 crystal structures covering 79,788 unique compositions. Of those structures, "77 571 (65%) have no shared crystallographic sites and are considered as ordered, whereas 41 113 (35%) are disordered."

Note what that base rate says before any model runs. A third of the experimentally characterised inorganic record is already disordered. This is not a rare pathology that computation happens to miss. It is a third of the field.

Their best model is a recurrent network that pools learned element embeddings, and it reaches 90% accuracy on held-out data with a ROC-AUC of 0.96. The comparison that gives those numbers meaning is one the authors provide themselves: a trivial classifier that called everything ordered would score 65%, because that is the class balance. Ninety against sixty-five is a real signal.

Now the design decision that governs everything downstream. The classifier reads composition. It takes the list of elements and answers whether that combination tends, in the experimental record, to be found disordered. It does not examine the specific predicted structure. This is what makes it cheap enough to sweep a database of millions, and it is also the exact boundary of what it can tell you, which I will come back to.

Two numbers

They pointed it at the Materials Project, the long-running computed database, and at GNoME, the Google DeepMind effort that announced 2.2 million predicted crystals and which this publication has been through before.

For the Materials Project: "we predict that 39–47% of the configurations in the MP database are likely to display substitutional disorder (excluding the overlap with the ICSD)," using conservative and balanced thresholds respectively. The authors flag the significance immediately, since the MP was seeded from the ordered subset of the ICSD: it has "accumulated a sizable fraction of potentially disordered systems through computational predictions."

For GNoME: "the classifier predicts a much higher fraction of disordered materials, ranging from 80% to 84%."

That second number is the one that went around the world. The institutional release said "over 80% of materials predicted by simulations showed signs of disorder." Phys.org reported that "in one instance, more than 80% of the proposed materials were affected." Neither piece I read named which database, and neither carried the 39 to 47 percent figure at all.

The sentence that did not travel

Here is the paragraph the 80% comes from, in the authors' own words:

the disjoint nature of the respective sets of compositions also represents a sizable covariate shift with respect to the training set of the classifier, so that our predictions here are attached with a significant degree of uncertainty

Unpack that, because it is the whole point. GNoME's value is that it went looking in chemical territory nobody had catalogued. The classifier learned what disorder looks like from the ICSD, which is the catalogue of things people have already made. So the database with the most headline-worthy number is precisely the one furthest from the model's experience, and the authors say so, attached to that number, in the same breath.

This is the pattern this publication keeps running into and it is worth naming plainly. The figure that escapes a paper is very often the one with the widest error bar, because breadth of claim and strength of evidence trade against each other, and press releases select on the first. It is the same shape as the machine-learned force fields that topped a benchmark and then could not finish a simulation, and the same shape as models that beat everything except a straight line.

What I will not do is flip it into a debunk, because the authors do not support that either. Their next sentence is "Nevertheless, even the conservative estimate of 80% is much higher than the ICSD and MP values," and they give a mechanism rather than a shrug: GNoME contains a much larger share of compositions with four or more elements, and compositions like that have, in their phrase, "a proclivity for disorder." More elements, more chances that two of them are similar enough to share a site. The direction of the finding is reasoned and probably right. The precision of "more than 80%" is what the paper declines to stand behind, and the coverage supplied anyway.

The caveat that went even further unread

There is a second limitation in the same paragraph, and it points the other way, which is why I find it the more interesting one:

the classifier currently does not consider positional disorder, so that the overall degree of disorder may also be underestimated

Positional disorder is the other family: a site occupied only part of the time, or a group that can sit in several orientations. It is not rare either. In the curated ICSD the authors "find more substitutionally (73%) than positionally (46%) disordered materials," so the phenomenon this classifier is blind to shows up in nearly half the disordered record.

So the honest summary of the paper is not "80% of AI predictions are wrong." It is narrower in one direction and wider in the other: for one kind of disorder, in a chemical space the model was not trained on, the estimate is large and uncertain, and it omits a second kind of disorder that is nearly as common. That is a more useful sentence than the headline and it is not much longer.

Why a spectroscopist reads this the way I do

My own research is microwave spectroscopy, which is a different corner of the field from solid-state crystallography, and I want to be careful not to borrow authority I do not have here. What transfers is not the technique. It is the habit.

A structure is not an observation. It is a model that was fitted to data, and every structure carries assumptions that were made to get the fit to close. Spend enough time doing that and you develop a reflex about which assumptions are load-bearing, because they are the ones that fail quietly: the fit still converges, the numbers still look reasonable, and the thing is wrong in a way the residuals do not shout about. "Every site holds one element" is exactly that species of assumption. It makes the calculation possible, it is true often enough to be invisible, and when it fails the output remains perfectly well-formed.

Which is also why the person who finds out is usually the one at the furnace, not the one at the screen. That gap between a database row and a substance you can hold is the recurring subject here, whether the row came from an autonomous lab or a generative model.

What the paper is actually for

It would be easy to file this as another AI-materials takedown, and that is not what the authors wrote. They built a filter to be used, not a verdict to be quoted. Konstantin Jakob, the first author, describes the purpose as steering: the models "predict whether a crystal is affected by disorder and steer material discovery towards computationally well-represented areas." Johannes Margraf, the senior author, puts the same thing as a condition: disorder "can be a critical stumbling block in computational materials science if it is not accounted for," and with these tools it "can be detected even in large-scale workflows and treated with suitable computational methods."

That is a modest, correct framing, and the paper is open access, so anyone screening a database can go and apply it. There are established computational methods for disordered systems. The failure this addresses is not that they do not exist, it is that a pipeline generating millions of candidates has no cheap way to know which candidates need them. A composition-level flag is a reasonable answer to that.

What it does not say

Two things, both of which follow from the classifier reading composition rather than structure.

It does not identify any specific entry as wrong. It says a composition belongs to a family that the experimental record shows disordered, which is a statement about a chemical neighbourhood, not a refutation of a particular row. Nobody has gone into a furnace and checked these.

And it inherits the ICSD's shape, including its gaps. The authors are explicit that their models are limited by the scarcity of data on disordered materials beyond that database, and specific about where this bites: on actinoids, they note the elements are "severely underrepresented in the ICSD however, so that this claim cannot be made with high confidence." A model trained on what people have made will be least reliable about what people have not made, which is the same structural problem as the covariate shift, showing up in a different place.

What I take from it is smaller than the headline and holds up better. Roughly a third of the experimentally known inorganic record is disordered, and computational discovery pipelines currently produce ordered structures only, so some meaningful fraction of every large predicted database describes an idealisation of a real material rather than the material. Now there is a cheap way to guess which fraction, and its authors were careful to say how much they trust it. The number worth remembering from this paper is not 80. It is 35.

Sources

  1. K. S. Jakob, A. Walsh, K. Reuter and J. T. Margraf, "Learning Crystallographic Disorder: Bridging Prediction and Experiment in Materials Discovery," Advanced Materials, published 23 October 2025, DOI 10.1002/adma.202514226, PMID 41128259, PMCID PMC12822528. Open access. (Primary source, opened and read via the PubMed Central copy. Sole source of every figure in this report: the definition of substitutional disorder; the curated ICSD set of 118,684 structures and 79,788 compositions with the 65% / 35% ordered-disordered split; the 73% substitutional and 46% positional figures; the RNN architecture, 90% accuracy, 0.96 ROC-AUC and the 65% trivial baseline; the 39–47% Materials Project estimate with its threshold note; the 80% to 84% GNoME estimate; and the four-or-more-element explanation. All five block and inline quotations, including the covariate-shift sentence, the positional-disorder sentence, the "Nevertheless" sentence, the "proclivity for disorder" phrase and the actinoid caveat, are verbatim from this paper. A preprint of the same work was posted to ChemRxiv on 24 July 2025 under an identical title, DOI 10.26434/chemrxiv-2025-f52qs; I could not open it, so no comparison between the preprint and the published version is made here.)
  2. Fritz Haber Institute of the Max Planck Society, "New International Study Uncovers Major Limitations in AI-Driven Materials Discovery," 11 December 2025. (Opened and read. Source of both researcher quotations, from Konstantin Jakob and Johannes T. Margraf, quoted verbatim, and of the "over 80%" framing characterised in this report. This release does not name the Materials Project or GNoME and does not carry the 39–47% figure.)
  3. "Limitations of AI-based material prediction: Crystallographic disorder represents a stumbling block," Phys.org, 9 December 2025. (Opened and read, solely to check what the coverage carried. Its "in one instance, more than 80% of the proposed materials were affected" is quoted verbatim. It names neither database, gives no second figure, and carries no uncertainty caveat. My characterisation of the coverage rests on this piece and the institutional release above; I did not survey every outlet that ran the story.)
  4. Prior reporting in this publication: Report 009, on GNoME's 2.2 million predicted crystals; Report 130, on benchmark scores that did not predict stability; Report 069, on models and linear baselines; Report 051, on the autonomous lab. (Context only. No claim in this report rests on them.)
Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports