← The Signal Report Work with me

Report 075 · AI in the Lab

Where AlphaFold quietly breaks the chemistry

A lab made the wrong protein by accident. When they measured it, it had twice the helix of the original and came in threes instead of ones. Four AI structure predictors, asked about that same sequence, handed back the original. Not garbage: the original, to within an ångström, with charged atoms parked in the greasy interior where charge is not allowed to sit. The interesting part is why the wrong answer looked so right, and how cheap it turns out to be to catch.

In microwave spectroscopy you never see a molecule. You measure a set of rotational constants, which are numbers describing how a molecule tumbles, and then a structure is whatever geometry reproduces those numbers. Get the geometry wrong and the numbers do not come out. The habit that grows out of years of that is simple and hard to unlearn: a structure is a proposal until a physical observable has had a chance to refuse it.

Predicted protein structures are proposals with unusually good manners. They arrive rendered, colored, ribboned, and confident. Nothing in the file tells you which parts were constrained by anything.

In late July 2026, Rensselaer Polytechnic Institute put out a release that got picked up as: AI protein-folding tools produce chemically impossible structures and need human oversight. That headline is accurate, and it is also easy to read the wrong way, as though the models emit obvious junk you would catch at a glance. The paper underneath says something considerably more uncomfortable. The junk looks fine.

The protein nobody meant to make

The paper is single-author: George I. Makhatadze of Rensselaer, in PNAS, published online 17 July 2026 and in the 21 July issue, volume 123, issue 29, article e2609610123. Its first sentence is one of the better opening lines in structural biology this year, and I am quoting it exactly:

A variant of the U1A protein containing four substitutions to ionizable residues was generated serendipitously due to a miscommunication.

Somebody in a lab made the wrong protein. That happens constantly. What happens less often is that somebody then measures the mistake carefully instead of throwing it out, and the mistake turns out to be interesting. Verbatim again: "Biophysical measurements reveal this variant has twice the helical structure of wild-type U1A and is trimeric, unlike the monomeric wild type."

Twice the helix. Three copies associating where the original sits alone. Four substitutions did that. Whatever this molecule is, it is not the parent protein wearing four new side chains. It is a different thing.

What four predictors said about it

Then the sequence went into the tools. Two from the deep-learning-with-alignments family, AlphaFold2 and RoseTTAFold2, and two protein-language-model tools, OmegaFold and ESMFold. All four returned structures, in the paper's words, "nearly identical to the wild-type (backbone RMSD < 1 Å)."

Sub-ångström backbone agreement is the language structural biologists use for the same structure. Four tools, two different architectural lineages, no communication between them, and every one of them confidently answered a question that had not been asked: what does the original protein look like.

Be precise about what kind of error that is, because it is not the kind people picture. There was no mangled chain, no knot, no chain running through itself, nothing that would make a reviewer squint. The models produced a clean, canonical, entirely presentable structure. It was just the structure of a different molecule.

The part chemistry does not allow

And then the specific violation, verbatim: "Surprisingly, these models predict ionizable residues buried within the nonpolar core, contradicting established physico-chemical principles."

This is the sentence to slow down on. Ionizable residues are the ones that carry or can carry a formal charge: aspartate, glutamate, lysine, arginine, histidine. Out in solvent, a charge is comfortable, because water molecules orient themselves around it and stabilize it. That stabilization is worth a great deal of free energy.

Now push that charged group into a protein's interior. The interior is hydrocarbon. There is no water there, the dielectric constant is low, and there is nothing to pay the charge back for the solvation it just gave up. The cost is called the desolvation penalty and it is not a rounding error.

Real proteins do bury charges, and when they do it, they pay: a salt bridge with an opposite charge, a hydrogen-bond network, a coordinated metal, a catalytic residue that has to be there for the chemistry to work. Buried, unpaired, unexplained charge is the structural equivalent of a line item with no receipt.

So the models did not merely produce an imprecise structure. They asserted a specific claim about where charge sits, and physical chemistry has held a firm opinion about that claim since long before any of these networks were trained.

It was not one strange protein

The obvious first response is that U1A is a quirky case and four substitutions is a stress test nobody promised to pass. So the paper scaled it up. Sequences were generated containing up to all twelve residues that make up U1A's nonpolar core. Verbatim:

Across thousands of sequences, and depending on the AI model used, the majority of predicted structures contained fully buried ionizable residues while still maintaining the overall U1A fold.

Then two more proteins of comparable size: acylphosphatase, a natural one, and TOP7, which is a famous de novo designed fold. Same result, described as "the same phenomenon."

The clause carrying the weight is "while still maintaining the overall U1A fold." The models had a fold they were sure of. When handed residues that do not belong in it, they kept the fold and put the residues wherever the geometry required, including into places the chemistry forbids. The confidence was in the shape, and the shape survived contact with facts that should have destroyed it.

The check costs an afternoon

Here is the part I would tape to a monitor. Verbatim: "However, short (50 ns) molecular dynamics simulations with physics-based force fields (CHARMM/AMBER) rapidly relaxed these structures, exposing the ionizable residues."

Fifty nanoseconds is a short simulation. For a protein this size on ordinary GPU hardware, that is hours, not a grant cycle. And the paper's recommendation follows directly: include "brief molecular dynamics simulations as a vital validation step for AI-generated structures."

Notice what that recommendation is not. It is not a bigger model, a better training set, or another network to check the first network. It is an independent instrument. The electrostatics in a classical force field were parameterized decades ago against measurements these predictors never saw, by people arguing about entirely different questions. When two methods that share no assumptions disagree, you have learned something. When two methods trained on the same data agree, you have learned considerably less. I made the same argument about asking three models the same question, and it is the same principle here from the other direction.

The strongest objection, taken seriously

A fair skeptic pushes back like this. Force fields have priors too. CHARMM and AMBER model charge with explicit Coulomb terms in explicit water, so of course a simulation built that way will drive a buried charge to the surface. Two methods encoding different assumptions disagreed. Why does the simulation get to be the referee?

That is a real objection and if simulation were the only evidence I would hold this loosely. It is not the only evidence, and the tiebreaker is not in either computer. It is on the bench. The variant was measured. Twice the helical content. Trimeric instead of monomeric. Those are experimental observables about a physical sample. The predictors said the molecule was unchanged, and the instrument said it was not.

The molecular dynamics is therefore not adjudicating a philosophical tie. It is the cheap proxy that flagged in a few hours what the wet-lab work confirmed. That is exactly what you want from a screening step, and it is why "run a short simulation" is a better recommendation than "be careful."

Why the tools are genuinely good and this still happens

None of this is a knock on the headline performance of structure prediction, and the paper is explicit about that: "while AI-based tools perform exceptionally on natural sequences."

Sit with the word natural. Natural sequences are sequences evolution already filtered. Every one of them folds. Every one of them is soluble enough to have survived in a cell. Every one of them has already settled its energetic accounts, including the ones about buried charge, because the alternatives were selected out long ago. The training data and the benchmark are drawn from the same pre-filtered population, and the filter did a lot of the work the model is being credited with.

The moment you type in a sequence that does not exist in nature, which is to say an engineered mutant, a designed protein, a rationally introduced substitution, a disease variant, you have stepped outside that population. Nothing marks the boundary. The output arrives with identical formatting and identical visual authority on both sides of it.

This is the same failure shape I wrote about when foundation models lost to deliberately simple baselines, though it runs the opposite way. There, the models barely changed their output when the input changed, and a benchmark that never checked for movement let it pass. Here the input changed enormously and the output did not move at all. Both are the same underlying question: was this system ever tested on the thing you are now using it for?

The accident is the methodological point

Come back to the miscommunication, because it is not a charming detail, it is the argument.

The best adversarial test case in this paper was not designed. Nobody sat down and reasoned their way to four substitutions that would fool four independent predictors while producing a genuinely different molecule. A mistake produced it for free, and it worked on the first try.

Which tells you something about how hard it is to imagine these failures deliberately. We build test sets out of the cases we can think of, and the cases we can think of are constrained by the same intuitions the models were trained on. If a random error in lab communication lands on a counterexample immediately, counterexamples are probably not rare. They are unlooked-for, which is a different thing, and a worse one, because unlooked-for failures do not show up in the benchmark table.

What I could not verify

Stated plainly, because the alternative is to let a claim ride that I did not check.

The RPI release, carried by Phys.org on 27 July and Technology Networks on 28 July, contains this line: "every tool tested rated its own accuracy higher than the results warranted." That is a claim about model confidence scores, and it is the most quotable sentence in the coverage. I could not confirm the numbers behind it. The PNAS full text is paywalled and its PMC deposit is embargoed until 17 January 2027; I read the abstract verbatim from the PubMed record and have quoted only from that. A preprint of related work exists on bioRxiv but rate-limited every attempt I made to retrieve it, so nothing here rests on it either.

Treat the confidence claim as the university's characterization of its own researcher's work, correctly attributed and plausibly true, and not as something I independently checked. Every quoted result above, the sub-ångström agreement, the buried residues across thousands of sequences, the two additional proteins, the 50 ns relaxation, is the paper's own language.

Two more limits worth saying out loud. This is one author, one lab, three proteins. And no rebuttal has appeared that I could find, which after three weeks means very little either way. A single paper is a single paper, and the correct response to one is interest, not conversion.

The author's own framing, from the release, is more modest than the headlines built on it:

AlphaFold is considered the gospel of the field. It is very good, and it does many things well. But occasionally it makes mistakes, because there simply isn't enough of the right kind of data in the model yet.

And his summary of the paper, which is four words long: "trust but verify."

Three questions this hands you

Is my sequence natural? If it exists in an organism, you are inside the population these tools were validated on. If you mutated it, designed it, or pulled it from a variant database, you are not, and the interface will not tell you. That single question sorts most of the risk.

Did the prediction move when the input moved? A model that returns essentially the parent structure for a substantially altered sequence has told you about its prior, not about your molecule. Sub-ångström agreement between a wild type and a multi-substitution variant is a warning, not a reassurance.

What independent instrument could contradict this? Not another model. Something with different assumptions: a force field, a spectrum, a gel, a melting curve. If nothing in your pipeline is capable of disagreeing with the prediction, your pipeline cannot detect that it is wrong.

The signal

The lesson here is not that AI structure prediction is unreliable. It is very reliable at the job it was measured on, and that job is predicting the structures of proteins that already exist.

The problem is that the job people actually want done, increasingly, is the other one: tell me what this protein will look like when I change it. That is the entire premise of protein engineering, of variant interpretation, of design. It is a different question, drawn from a different population, and it was never the one on the benchmark.

The molecule in this paper announced the gap loudly, by existing and being visibly different. Most of the time, the gap will not announce itself. It will hand you a beautiful ribbon diagram of a protein you did not ask about, and the only thing standing between you and believing it is whether you bothered to run something that could have said no.

Sources

  1. George I. Makhatadze, "The accuracy of electrostatic interactions captured by AI protein structure prediction models," Proceedings of the National Academy of Sciences, 21 July 2026, vol. 123, issue 29, article e2609610123 (published online 17 July 2026), DOI 10.1073/pnas.2609610123, PMID 42467514. (Primary, peer reviewed. Source of every verbatim quotation in this report: the "serendipitously due to a miscommunication" origin, the twice-the-helix and trimeric biophysical measurements, the four named predictors, the "nearly identical to the wild-type (backbone RMSD < 1 Å)" result, the buried-ionizable-residue finding, the thousands-of-sequences result, acylphosphatase and TOP7, the 50 ns CHARMM/AMBER relaxation, the "perform exceptionally on natural sequences" concession, and the recommendation to add brief molecular dynamics as a validation step. Access limit, stated for the record: the full text is paywalled and the PMC deposit (PMC13389394) is embargoed until 17 January 2027. I read the complete abstract verbatim from the PubMed record and have quoted only from it. I did not read the figures, methods, or any confidence-score data.)
  2. "AI tools for predicting protein folding produce chemically impossible structures and need human oversight," Phys.org, 27 July 2026, based on a news release from Rensselaer Polytechnic Institute. (Coverage, opened and read. Source of the Makhatadze quotations used above, and of the "every tool tested rated its own accuracy higher than the results warranted" claim, which is flagged in the report as unverified against the paper itself.)
  3. "Blind Spots Revealed in AI Tools for Protein Structure Prediction," Technology Networks, 28 July 2026, republished from the same Rensselaer release. (Coverage, opened and read independently to confirm the Makhatadze quotations appear identically in both renderings of the release rather than in one outlet's paraphrase.)
Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports