← The Signal Report Work with me

Report 176 · AI in the Lab

What 37,000 AI agents actually proposed

A Stanford team built a "virtual biotech" out of AI agents, and the coverage said the swarm found a lung-cancer drug candidate. What it proposed was a B7-H3 antibody-drug conjugate. A B7-H3 antibody-drug conjugate has been in human trials since November 2019. The genuinely new piece is narrower, it is interesting, and nobody has tested it at a bench yet.

The paper is real and worth reading. It appeared in Science on 17 September 2026, from James Zou's group at Stanford with a co-author at PHD Biosciences, after a preprint went up on bioRxiv in February. I read the preprint in full and the published abstract. The published body may differ in places, and I flag below where that could matter.

My question is the one this beat always asks: what did the model do, and what did the people and the prior literature do? Here, the answer changes the headline.

What the system is

The "virtual biotech" is an organization chart made of language-model agents. A Chief Scientific Officer agent takes a question, hands pieces of it to specialist agents (statistical genetics, single-cell biology, pharmacology, clinical trials and so on), and stitches their reports together. A reviewer agent checks the specialists' work. Per the preprint, the agents run on Claude Sonnet 4.5, with Claude Haiku 4.5 for two support roles (a chief of staff and the reviewer), and they reach public datasets through tool connections: GWAS catalogs, single-cell atlases, ChEMBL, ClinicalTrials.gov and others.

The paper shows it at three decision points. The famous number, 37,000 agents, belongs to the first one, not the drug.

Demo one: 37,000 agents reading trial records

The 37,000 were mostly clerks. Thousands of "clinical trialist" agents read 55,984 trial records and coded each one's outcome, then the system linked each drug's target to gene-expression data. The finding: drugs aimed at genes expressed in only a few cell types were more likely to advance. The abstract says 48% more likely to reach market, with 32% fewer adverse events; the preprint adds 40% more likely to get from Phase I to Phase II.

Two things make this number smaller than it sounds. First, the coding was checked against people, and it was good but not perfect: on 100 randomly chosen trials, the agent agreed with a human reviewer on 89.7% of primary endpoints (78 of 87), 83.9% of secondary endpoints and 92.4% of adverse-event rates. Second, the authors say it themselves: "our trial analyses are observational and should be interpreted as such." They adjusted for phase, year, therapeutic area and modality and ran permutation tests, which is the right work. It is still a correlation across trials, not a rule for picking targets.

Demo two: the lung-cancer "drug"

This is the part that became the headline. The team asked the system to evaluate B7-H3 (also called CD276), an immune-dampening protein, as a lung-cancer target. The genetics agent found no germline association. The single-cell and spatial agents found that, in tumors, B7-H3 was strongly expressed in cancer-associated fibroblasts, the support cells around a tumor, and that high-B7-H3 regions had fewer immune cells nearby. Small molecules were ruled out because no druggable pocket turned up. The system proposed an antibody-drug conjugate: an antibody that finds B7-H3 and carries a toxic payload to it. The preprint puts the cost of that case study at $46 in API credits.

Note what was produced: a strategy, written up. No molecule was made. No cell was dosed. The Science abstract describes it accurately, as integrating evidence "to propose a therapeutic strategy in lung cancer." The coverage was less careful. Business Today's headline asked, "AI biotech without humans?"

The validation that was already in the clinic

The preprint's evidence that the strategy was good is a real drug. In August 2025, the FDA granted Breakthrough Therapy Designation to ifinatamab deruxtecan, a B7-H3-targeted antibody-drug conjugate from Daiichi Sankyo and Merck, for previously treated extensive-stage small cell lung cancer. The preprint calls it "a first-in-class B7-H3–targeted ADC that demonstrated a 48.2% objective response rate in a phase II trial." The agents ran without web access, on a model with a stated knowledge cutoff of January 2025, so the authors argue the trial's results provide "independent external support for the system's mechanism-driven therapeutic design."

Here is the problem with that logic. The designation came after the cutoff. The drug did not. ClinicalTrials.gov lists Daiichi Sankyo's first-in-human study of ifinatamab deruxtecan (code DS-7300a), a Phase 1/2 trial in advanced solid tumors, with a start date of 3 November 2019. That is more than five years before the model's cutoff. The idea of hitting B7-H3 with antibodies was not obscure either: the preprint itself cites a 2021 review titled "B7-H3: an attractive target for antibody-based immunotherapy," and writes that existing literature motivated "antibody- and ADC-based programs."

So the "independent" match is between the system and a drug class that had been public, and in patients, for years. I cannot know what was in the model's training data, and neither can the authors from the outside. But a language model trained through January 2025 proposing a B7-H3 ADC is not the same event as a human team arriving there blind. A regulatory milestone that postdates the cutoff does not make the underlying idea postdate it.

One more detail is worth noticing. The preprint's account of the system's competitive-landscape search says it found the approved HER2 ADC trastuzumab deruxtecan as a modality precedent, found the B7-H3 antibody enoblituzumab in Phase 2 for prostate cancer, and concluded that no B7-H3 therapy had been approved. All true. The account does not mention the B7-H3 ADC already in trials. For a system whose job is to map the competitive landscape before a company commits money, that is the item it most needed to find.

What is genuinely new

The authors are clear about where they think the originality lies, and it is not "ADC against B7-H3." It is the fibroblast angle: that in lung cancers B7-H3 is "more strongly overexpressed in cancer-associated fibroblasts, spatially associates with immune exclusion, and stratifies poor clinical outcomes," which they call "an interpretation that is not directly identified from any single dataset or previous publication." Prior work, they write, looked at B7-H3 "largely through a cancer-cell–centric lens."

That is a real, testable hypothesis, and pulling it out of genetics, single-cell, spatial and outcome data in one pass for $46 is a genuine productivity result. It is also, so far, computational. It came from public datasets, it has not been checked in tissue or cells, and the drug that "validated" the strategy was not designed around fibroblasts. The Stanford team appears to know this: Business Today reports they stressed that the findings still need physical experiments and, ultimately, human trials.

The third demo, briefly

The system also reread a terminated Phase II ulcerative-colitis trial of vixarelimab, an antibody against OSMR, and argued it should have enrolled patients by OSMR expression rather than all comers. It is a sensible post-mortem built on public expression data from other trials. It is also hindsight, and it has not been tested prospectively.

How to read the next "AI agents discovered a drug" story

Molecule, or memo? Ask whether anything was synthesized or tested. Here, the output was a proposal.

Is the "validation" older than it looks? When a paper says a later event confirms its AI's idea, check when the idea itself first went public. A trial registry date takes thirty seconds to look up.

Where is the 37,000? Big agent counts tend to belong to the clerical part of the work. That part can be useful, and it should come with an error rate, as this one did.

What do the authors claim, versus the headline? The Science abstract says "human-guided" and "inform." That is the honest version, and it is the version to repeat.

I use these tools daily, and I think systems like this will save labs real time on exactly the literature-and-database synthesis shown here. That is the reason to describe them precisely. An AI that reproduces a drug strategy already in Phase 2 is a decent sanity check. An AI that found a lung-cancer drug would be news. This paper is the first thing.

Disclosure, plainly: Claude, the model family the virtual biotech runs on, is one of the tools I use in my own work, including research and drafting for this publication. I am a customer of Anthropic's products and have no other relationship with the company, Stanford, PHD Biosciences, Daiichi Sankyo or Merck; none of them asked for, saw or reviewed this report. Nothing here is sponsored and no link earns a commission; here's the full policy.

Sources

  1. H. G. Zhang, P. Eckmann, J. Miao, A. B. Mahon and J. Zou, "The Virtual Biotech: A multi-agent AI framework for therapeutic discovery and development," Science, published 17 September 2026, article eaeg6779, DOI 10.1126/science.aeg6779, PMID 42752167. (Primary source. Abstract read via Crossref and Europe PMC; the published full text was not read. Source of the 37,000 agents, 55,984 trials, 48% / 32% figures, and the quoted phrases "propose a therapeutic strategy in lung cancer," "human-guided" and "inform.")
  2. H. G. Zhang, P. Eckmann, J. Miao, A. B. Mahon and J. Zou, "The Virtual Biotech: A Multi-Agent AI Framework for Therapeutic Discovery and Development," bioRxiv preprint v1, posted 23 February 2026, DOI 10.64898/2026.02.23.707551, CC BY-ND. (Primary source, the 45-page PDF opened and read in full. Source of: the agent architecture and models (Claude Sonnet 4.5, Claude Haiku 4.5); the 40% Phase I-to-II figure; the 100-trial human check (89.7%, 78/87; 83.9%; 92.4%, 61/66); the "observational" quote and adjustments; the B7-H3 genetics, fibroblast and spatial findings; the ADC proposal and $46.00 cost; the no-web-access condition and January 2025 cutoff; the competitive-landscape search naming trastuzumab deruxtecan and enoblituzumab; the quoted ifinatamab deruxtecan, "independent external support," "cancer-cell–centric lens," fibroblast and "antibody- and ADC-based programs" passages; the citation of the 2021 Kontos et al. review; and the vixarelimab / OSMR case. Any changes made between this preprint and the Science version are not reflected here.)
  3. ClinicalTrials.gov, NCT04145622, "Study of Ifinatamab Deruxtecan (DS-7300a, I-DXd) in Participants With Advanced Solid Malignant Tumors", sponsor Daiichi Sankyo, Phase 1/Phase 2. (Primary registry record, read via the ClinicalTrials.gov v2 API. Source of the 3 November 2019 start date.)
  4. Merck, "Ifinatamab Deruxtecan Granted Breakthrough Therapy Designation by U.S. FDA for Patients with Pretreated Extensive-Stage Small Cell Lung Cancer", news release, 18 August 2025. (Opened and read. Source of the designation, date and indication. The release does not state the 48.2% response rate; that figure is quoted from the preprint, which attributes it to the Phase II IDeate-Lung01 trial.)
  5. Edd Gent, "Virtual Biotech Company Puts 37,000 AI Agents to Work on Drug Discovery", Singularity Hub, 18 September 2026. (Coverage, opened and read. A careful write-up that does note the strategy matched an existing drug; consulted for context only.)
  6. Business Today Desk, "AI biotech without humans? Stanford deploys 37,000 agents to hunt new therapies", Business Today, 20 September 2026. (Coverage, opened and read. Source of the quoted headline and of the report that the researchers stressed the need for physical experiments and human trials.)
  7. Prior reporting in this publication: Report 165, what Claude found in phage DNA; Report 022, AI drugs in Phase 1 versus Phase 2. (Context only. No claim in this report rests on them.)
Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports