← The Signal Report Work with me

Report 152 · AI in the Lab

The benchmark nobody outside can run

On 14 September a consortium of five drug companies reported that fine-tuning an open protein-folding model on more than twenty thousand of their private crystal structures beat every public model tested. The result looks real and the write-up is unusually candid about its own limits. It is also, by construction, a number that nobody outside those five companies can check, reproduce, or dispute. That is worth understanding properly, because a lot more science is about to arrive in this shape.

Here is the story as it was reported. The AI Structural Biology Network, a consortium made up of AbbVie, Astex Pharmaceuticals, Bristol Myers Squibb, Johnson & Johnson and Takeda, working with Apheris and Mohammed AlQuraishi's lab at Columbia, took OpenFold3 Preview 2 (an open-source model in the AlphaFold 3 family) and fine-tuned it across 20,167 protein-ligand structures held privately inside those five companies. The structures never left their owners. Each company trained locally, and only model parameters were sent to a central aggregator.

The resulting model, AISB-1-Fed, was then compared against the base model and several public ones. The consortium's own summary of the headline numbers, on 1,056 held-out structures:

  • Fraction of structures with protein-ligand interface lDDT at or above 0.8: 35.6% for the base OpenFold3 Preview 2, 52.1% for AISB-1-Fed, against 40.9% for Boltz-2.
  • Fraction with ligand pose error (bisyRMSD) at or below 2 angstroms: 28.9% for the base model, 46.8% for AISB-1-Fed, against 36.5% for Boltz-2.
  • On a 443-structure subset with protein-protein interface annotations, 62.6% for AISB-1-Fed against 40.9% for Boltz-2, despite protein-protein quality never being an optimisation target.

Nature's news desk covered it the same day. AlQuraishi's line there is a fair summary of the finding: "You add all this data, and you get a pretty big bump in performance."

I want to be clear before going further: I think this is good work, and the consortium's write-up is more forthcoming than most peer-reviewed papers I read. They published their evaluation protocol, their sample-selection rule, their distributions, bootstrapped confidence intervals, a public-data-only control model, a privacy attack assessment, and a limitations section that names the central weakness in a bolded heading. Very little below is a criticism of what they wrote. Most of it is a criticism of how a result in this shape travels once it leaves them.

What "held-out" means here

In the ordinary reading, a held-out test set is the thing that makes a benchmark a benchmark. Somebody else can take your model, run it on the same set, and get your number. That is the whole mechanism. The number is not really the claim: the reproducibility is the claim, and the number is a summary of it.

This test set works differently. Each partner held out 5% of its own private structures. The split was made at the project level, so all structures from one medicinal-chemistry programme went entirely to training or entirely to evaluation, and partners were additionally asked to choose evaluation projects that were unlikely to appear in another partner's training data. Those are careful choices, and they are the right ones. They are also an admission: a naive random split of structures from dense medicinal-chemistry series would have leaked badly, because a series is dozens of near-identical molecules on one target.

The consortium states the residual problem itself, under a heading of its own:

A partner-held-out benchmark, not a fully external test: the evaluation structures come from the same five partners that contributed training data, so this is a partner-held-out benchmark rather than a fully external test of protein-ligand co-folding.

And then the part I found most interesting, because it is a genuine structural limit rather than a resource problem:

measuring full cross-partner similarity is itself hard under federation, since the structures never leave their owners, which points to an opportunity for better privacy-preserving federated metrics.

Read that carefully. The same privacy architecture that makes the training possible also makes it impossible to measure how similar the test set is to the training set across companies. Nobody is hiding the leakage number. The leakage number cannot currently be computed by anyone, including them. That is an honest statement of an unsolved problem, and it did not survive into the coverage.

The comparison is run by one side

Boltz-2 and ProtenixV1 are public models. The test set is not. So the reference numbers in this comparison were produced by the consortium, running other people's models on data only the consortium holds, through a metric pipeline the consortium built. They say so plainly: all reference models were evaluated on the same 1,056 private structures with the same metric pipeline.

There is nothing improper about that. It is the only way the experiment could have been run at all. But it removes the adversarial step that normally does the quality control in benchmarking. Ordinarily, if you publish that your model beats mine, I can go and check, and the knowledge that I can check is most of what keeps the claim honest. Here the authors of Boltz-2 cannot rerun their own model on this set, cannot inspect how it was prepared for their model's input conventions, and cannot test whether a different confidence heuristic would have changed their ranking.

That last one matters more than it sounds. Every structure was predicted 25 times, five diffusion samples across five random seeds, and the reported result for each model is its single best prediction as chosen by that model's own native confidence score. For OpenFold3, AISB-1-Fed and ProtenixV1 that score is a composite of interface confidence, alignment confidence, and penalties for disorder and clashes. For Boltz-2 it is a weighted average of complex pLDDT and interface confidence. Those are different selectors. A model is therefore being scored partly on how well it ranks its own guesses, on a class of target it has never seen, and a model fine-tuned on that class has an obvious advantage at that ranking task that is separate from its advantage at folding. This is a standard and defensible protocol. It is also a place where a fine-tuned model can pick up points that the headline framing ("more accurate") does not distinguish.

The number that did not travel

Here is the part I would put in the first paragraph if I were writing the press release, and it is in the consortium's post but in none of the coverage I read.

The interface quality metric is bimodal. In their words: roughly 28% of per-structure values fall below 0.2 (failures) and another 28% above 0.9 (near-perfect), with the middle largely empty. The ligand pose metric behaves similarly, with a long tail: 29 to 41% of structures exceed 12 angstroms depending on the model.

Sit with that. A 12-angstrom ligand placement error is not a slightly imprecise pose. It is the drug in a different part of the protein. So the distribution is not a bell curve that shifted to the right. It is two piles, one of things the model basically nailed and one of things it got completely wrong, and the improvement moved structures from the second pile to the first while leaving a large failure pile in place.

The consortium explains exactly why they chose thresholded fractions over means: "Means land in regions of the distribution that few structures actually inhabit." That is a correct and rather elegant piece of statistical honesty, and it is the reason the reported metric is a percentage rather than an average error. But the same fact carries a second implication they do not spell out, and the coverage certainly did not: "52.1% of structures scored well" also means that on the order of a quarter of them failed outright, for the improved model, on the partners' own targets. In drug discovery that is a workable tool. It is not what "predicted more than half of them to a high level of accuracy" conveys to a general reader, which is a model that is usually right and sometimes a bit off.

Why the private data helps, which is the real finding

Strip away the benchmark question and there is a solid, believable result underneath, and it rests on an argument about data rather than about architecture.

The Protein Data Bank holds more than 200,000 experimentally determined structures and was the foundation of AlphaFold 2. But it is not full of drug-like molecules. The consortium's figures: since the original AlphaFold 3 data cutoff the PDB has grown by roughly 70,000 structures, almost all cofactors, metabolites, ions and crystallographic additives; only about 10,600 public structures contain an approved or investigational drug, and only about 3,000 of those are new since the cutoff. Paul Mortenson of Astex gave Nature a similar figure for drug-like co-structures, "maybe just 10,000." Against that baseline, 20,167 structures from live medicinal-chemistry programmes roughly triples the drug-relevant training data.

And crucially the consortium built a control for this. AISB-1-OF3p2-all-PDB is the same base model fine-tuned on all available public PDB with a later cutoff (19 November 2025) and no private structures. AISB-1-Fed beats that too, which is what isolates the private data as the cause rather than simply having newer public data. That control is the single most persuasive thing in the write-up, and it is the part I would defend if somebody called the whole result marketing.

Where I am coming from

My published research is in microwave spectroscopy, not structural biology. I have not run OpenFold3 and I have no stake in any of the companies named here.

What I do have is a decade of habits from electronic warfare, and this is a measurement-integrity problem in the exact shape that field trains you to recognise. In EW you learn quickly that a performance figure produced entirely inside one organisation, on that organisation's own scenario, against that organisation's model of the adversary, is a planning number and not a truth. It can be produced in complete good faith by careful people and still be wrong in ways nobody in the building is positioned to notice, because everyone shares the same assumptions about what the test should look like. The fix was never to distrust the people. It was to get a reference somebody else also holds.

That is precisely what the AISB result does not have and, under its current architecture, cannot have. Which is why the most important sentence in the whole announcement may be the one about needing better privacy-preserving federated metrics. It names the missing instrument.

I have written this beat's version of the same complaint before, from the other direction: Report 087 on what happens when a model's output is treated as an instrument reading, and Report 130 on a benchmark ranking that inverted when the models were made to do the job instead of the test. The recurring lesson is that a benchmark is a social arrangement more than a statistical one, and when you remove the part where somebody else can run it, what is left is a self-report.

What I could not confirm

There is no paper. The result is published as a company blog post. It has not been peer-reviewed, and Nature reports that the team plans to submit a paper to a peer-reviewed journal. Every number above comes from that post, which I opened and read in full, or from Nature's news coverage of it, which I also opened and read. Nature's own characterisation, quoted here because it is the cleanest statement of the situation: "The study, described in a blog post, has not been peer-reviewed, and the model is not publicly available."

I did not verify any measurement. I did not run OpenFold3, Boltz-2 or ProtenixV1, I do not have access to the 1,056 evaluation structures, and no independent party does. Nothing in this report is a claim that the consortium's numbers are wrong. They may well be exactly right. The point is narrower and, I think, more durable: they are currently unfalsifiable, and that is a property of the result rather than of the people who produced it.

Technical reports and preprints I did not open. The OpenFold3 Preview 2 technical report, the Boltz-2 preprint and the ProtenixV1 preprint are all linked from the consortium's post and I did not read them. So I am taking the descriptions of those models, including which confidence heuristic each exposes, from the consortium's account of them. Nature's article also cites a 2026 paper in Nature Structural & Molecular Biology for the claim that co-folding accuracy degrades sharply on molecules dissimilar to the training set; I could not locate and open that paper, so I have left the claim out rather than repeat it.

The privacy assessment is theirs. The consortium reports a formal privacy risk assessment covering data reconstruction and membership inference, and states that membership-inference true positive rates ran from roughly 0.05 to 0.32 at a zero false-positive rate under deliberately attacker-favourable conditions. I have not evaluated that assessment and it is outside my competence to referee. I mention it only because it sits in the same category as the accuracy claim: internally produced, externally uncheckable.

The signal

Federated training on pooled proprietary data is going to become common, in structural biology first and then anywhere a valuable dataset is legally unpoolable. Clinical records, manufacturing telemetry, materials characterisation, battery cycling logs. The privacy machinery that makes it possible has a side effect that nobody has solved: it also encrypts the evidence.

So when the next one of these lands, three questions separate a real result from a press release, and this one answers all three well:

Is there a public-data-only control trained the same way? Without it you cannot tell new private data from merely newer data. AISB built one, and it is why their claim stands up.

Did they publish the distribution, or only the headline fraction? A bimodal metric with a quarter of cases failing tells a very different operational story than the same percentage drawn from a smooth distribution. AISB published the distribution. Almost nobody reporting on it read it.

Do they say what they cannot measure? The sentence about cross-partner similarity being unmeasurable under federation is the most useful thing in the document, and a less careful group would simply have omitted it and nobody would have known.

The honest read is this. The consortium produced a strong, well-controlled internal result, documented it better than most published papers, and named its own central limitation in bold. Then it was written up as a model that beats AlphaFold-class systems, and the limitation did not make the trip. The failure mode in this beat is almost never that scientists lied. It is that they told the truth in paragraph nineteen.

Sources

  1. Avelino Javer, Nicolas Gautier, Benedict W. J. Irwin, Alwin Bucher, José-Tomás Prieto, Inken Hagestedt, Ruda Porto Filgueiras, Mees Hendriks, Lewis Mervin, Mark Sharpley, Ian Hales, Robin Röhm (Apheris); Sreeja Kutti Kandy, Frank Oellien, John Karanicolas (AbbVie); Maria Kadukova, Lucian Chan, Carl Poelking, Paul Mortenson, Chris Murray (Astex Pharmaceuticals); Matt Pokross, Veerabahu Shanmugasundaram, Payal Sheth (Bristol Myers Squibb); José Carlos Gómez Tamayo, Gary Tresadern (Johnson & Johnson); Edward King, Prashanth Vishwanath (Takeda); Lukas Jarosch, Vinay Swamy, Mohammed AlQuraishi (AlQuraishi Lab, Columbia University), "Federated Training Dramatically Improves the Accuracy of Protein-Ligand Co-folding on Private Pharma Structures," Apheris, published 14 September 2026. (PRIMARY, full post opened and read. NOT peer-reviewed; published as a company blog post, and the consortium states that both the underlying structures and the trained model weights remain private. Source for: the 20,167 private training structures and the 1,056 held-out evaluation structures; the identity of the five partner companies and the AlQuraishi Lab collaboration; the base model OpenFold3 Preview 2; the PL-lDDT ≥ 0.8 figures of 35.6%, 52.1% and 40.9% for Boltz-2; the bisyRMSD ≤ 2 Å figures of 28.9%, 46.8% and 36.5% for Boltz-2; the 62.6% versus 40.9% protein-protein interface result on the 443-structure annotated subset; the AISB-1-OF3p2-all-PDB public-only control and its 19 November 2025 PDB cutoff; the 5% per-partner project-level evaluation split and the instruction to partners to select evaluation projects unlikely to appear in another partner's training data; the 5 diffusion samples across 5 seeds giving 25 predictions per structure and the per-model native confidence selector, including the specific composition of those selectors; the statement that all reference models were evaluated on the same 1,056 private structures with the same metric pipeline; the bimodality of the PL-lDDT distribution at roughly 28% below 0.2 and 28% above 0.9, and the 29–41% of structures exceeding 12 Å on bisyRMSD; the rationale for thresholded fractions, quoted verbatim; the PDB growth figure of roughly 70,000 structures since the AlphaFold 3 cutoff and the roughly 10,600 public structures containing an approved or investigational drug with roughly 3,000 new since that cutoff; the membership-inference true positive rates of roughly 0.05 to 0.32 at zero false-positive rate; and both quoted passages from the Interpretation and limitations section.)
  2. "Drug firms' secret data supercharge AI protein models," Nature news, 14 September 2026 (clarification issued 15 September 2026), DOI 10.1038/d41586-026-02882-x. (Opened and read in full. Independent journalistic account of the same announcement. Source for: the statement that the study has not been peer-reviewed and the model is not publicly available, quoted verbatim; the Mohammed AlQuraishi quotation, verbatim; Paul Mortenson's estimate that the PDB holds "maybe just 10,000" experimentally determined structures with drug-like molecules; the note that the team plans to submit a paper to a peer-reviewed journal; and the existence of the separate OpenBind public-data project. The clarification of 15 September concerned AbbVie's location and does not affect anything quoted here.)
  3. Onur Oncer, "Your AI result won't reproduce," The Signal Report 087, and "The force field that passed the benchmark," The Signal Report 130. (Earlier reports in this beat on model outputs treated as instrument readings, and on a benchmark ranking that did not survive contact with the task it was supposed to predict.)

Scope note: this report analyses how a publicly announced but externally unverifiable machine-learning result should be read. It does not assert that any figure reported by the AI Structural Biology Network is incorrect, and it makes no claim about the scientific merit of OpenFold3, Boltz-2, ProtenixV1 or any product or service of the organisations named. No measurements or model runs were performed for this report. The author has no financial interest in, and no relationship with, any company or laboratory named here.

Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports