Report 146 · AI in the Lab
What the AI scientist actually found for AMD
An AI system called Robin proposed an existing glaucoma drug as a treatment for dry age-related macular degeneration, and the work is now a peer-reviewed paper in Nature. The result is real. It is also a fluorescence readout from cells in plastic wells, three wells per condition, and the paper says plainly that a disease model and a randomized trial are still ahead. Here is how to read it.
The headline version travels well: an AI scientist found a new treatment for a major cause of blindness. The paper behind it, "A multi-agent system for automating scientific discovery," by Ali Ghareeb, Benjamin Chang and colleagues at FutureHouse and Fordham University, appeared in Nature on 19 May 2026. It is open access, which means the gap between the headline and the evidence can be checked by anyone. I checked it.
My short verdict: this is a good paper about a workflow, and a modest, honest paper about a drug. Most coverage swapped those two.
What Robin is
Robin is not one model. It is a coordinator that calls three specialized agents. Two of them, Crow and Falcon, are literature agents built on PaperQA2 that write short and deep summaries of the published science. The third, Finch, analyses experimental data such as flow cytometry and RNA sequencing by writing and running analysis notebooks. According to Nature's press materials, the underlying language models were OpenAI's o4-mini and Anthropic's Claude 3.7.
Given a disease, Robin asked general questions about its pathology, identified ten candidate disease mechanisms, had an LLM judge rank proposed lab models of those mechanisms head to head, and settled on one: the retinal pigment epithelium, the cell layer that eats the worn-out tips of photoreceptors, doing that job poorly. Its strategy was to make those cells eat better. It then proposed 30 existing drugs, ranked them in another LLM-judged tournament, and handed over a list.
The authors report that in this workflow Robin analysed 551 papers in 30 minutes, against an estimated 294 hours for a person, and that a full run cost about US$10.76 in agent calls. That is the part of the paper I find most impressive, and it is almost entirely about speed of literature synthesis.
What the humans did
The paper's strongest claim is this sentence from the abstract: "All hypotheses, experimental directions, data analyses and data figures in the main text of this report were produced by Robin." Read the methods, though, and a list of human decisions sits underneath it, every one disclosed by the authors themselves.
The ranked drug list "was then reviewed by human scientists, and the top drug candidates were tested in the laboratory by executing a human-generated experimental protocol based on the assay suggested by Robin." The protocol was written by people.
Robin suggested feeding the cells real photoreceptor outer segments. The team used fluorescent pHrodo beads instead, "due to availability." Robin suggested primary or stem-cell-derived retinal cells. The team used the ARPE-19 cell line for the first screen "to expedite our evaluation of the hypotheses by Robin." For the RNA sequencing, "Read demultiplexing and alignment was performed by a human." The figures were "formatted for readability in publication by a human."
None of this is a scandal. It is how a real lab runs, and the authors call the approach "semi-autonomous" in their own abstract. But it matters for the headline, because the choices that most shape a cell-assay result (which cells, which substrate, which protocol) were human choices. Anyone who manages lab operations knows that swapping a cell line and a substrate is not a formatting detail. It changes what the number means.
What was actually measured
The assay is simple to describe. Treat retinal pigment epithelial cells with a drug for an hour, add pH-sensitive fluorescent beads, wait three hours, and count glowing cells by flow cytometry. The beads light up only in the acidic interior of the cell's digestive compartments, so more fluorescence means more beads were swallowed.
Round one tested the top five picks and found that Y-27632, a research-grade Rho kinase (ROCK) inhibitor, increased uptake. Robin then proposed RNA sequencing of treated cells, and Finch's analysis flagged a threefold rise in ABCA1, a gene for a lipid export pump, which the authors describe as a possible new target.
Round two tested ten more drugs. The headline number comes from here: ripasudil, "a ROCK inhibitor approved for treatment of glaucoma in Japan, outperformed Y-27632 and increased RPE cell phagocytosis 1.89-fold compared with dimethyl sulfoxide (DMSO) controls." The human re-analysis of the same data put it at 1.75-fold. The figure legend gives the sample size: three wells.
To their credit, the authors then went further. They repeated ripasudil in geriatric primary human retinal pigment epithelial cells, four wells per condition, and it again beat Y-27632. A cytotoxicity check found no sign that higher doses were killing cells. A second compound, KL001, a circadian clock modulator, also came up as a hit in those primary cells. And in a useful control, they gave the same prompt to OpenAI's Deep Research agent and screened its suggestions: "None of the suggested drugs by Deep Research were hits in this assay."
What "new treatment" skips
A cell swallowing more beads in a dish is a mechanism result. It is not a patient seeing better, and it is not even an animal retina clearing more debris. There is a chain of steps between the two, and the paper does not claim to have taken any of them. Its discussion is explicit: "This hypothesis would of course require validation in a suitable disease model and ultimately in a randomized, placebo-controlled trial to confirm clinical validity."
Elsewhere it adds that "in vivo validation of both drugs would be necessary for definitive comparison" of ripasudil against Y-27632. No animal work appears in the paper. That is exactly where most promising cell-culture drug repurposing stalls, which is the subject of an earlier report here on why AI-found drugs clear early stages and then stall.
There is a novelty question too. The paper itself notes that "the ability of ROCK inhibition to enhance phagocytosis in RPE cells is already known." The authors' claim is narrower and fair: that, to their knowledge, ripasudil had never been proposed for dry AMD. In other words, Robin connected a known cell effect to an approved drug and a disease. That is valuable, and it is the kind of connection a thorough human literature review can also make, just far more slowly. When The Scientist covered the preprint in June 2025, neuroscientist Konrad Kording of the University of Pennsylvania said reading it prompted him to search whether the strategy had already been explored, and put the larger ambition this way: "It's a big dream, and I'm not sure if the time has come for it yet." The same article reports the FutureHouse team acknowledging the result is not yet a "move-37" moment.
One disclosure worth reading
The competing-interests statement notes that most of the authors, including all three senior authors, "hold shares in Edison Scientific, a spin-out of FutureHouse." That does not make the data wrong, and it is disclosed exactly where it should be. It is a reason to read the word "discovery" in the framing with the same care you would give any company describing its own product. I hold the Signal Report to the same rule: where there is a stake, it gets named.
How I would read the next "AI discovers a drug" headline
Most of my career outside the lab has involved separating a demonstration from a capability, and my lab work is in microwave spectroscopy, not ophthalmology, so I am not judging the biology of retinal disease here. What I can offer is a checklist that works on any of these stories, and Robin's paper answers every item honestly if you go looking:
Where did the number come from? A cell line, primary cells, an animal, or people. Here: a cell line, then primary human cells, both in wells.
How many? Here: three wells, then four.
Who chose the assay conditions? Here: the AI proposed, humans chose, and humans changed two key conditions.
Was the mechanism already known? Here: yes, for the drug class. The new link is the specific drug and disease.
What do the authors say comes next? Here: a disease model and a randomized, placebo-controlled trial.
Run those five questions and the story sorts itself out. Robin is a genuinely useful literature and analysis engine that proposed a sensible, testable, cheap-to-check idea and helped a lab confirm it in a dish in weeks. That is worth a Nature paper. It is not yet a treatment for blindness, and the people who built it say so in the paper.
Sources
- Ali E. Ghareeb, Benjamin Chang, Ludovico Mitchener, et al. (corresponding authors Andrew D. White, Michaela M. Hinks and Samuel G. Rodriques), "A multi-agent system for automating scientific discovery," Nature 655, 497–505, published 19 May 2026, DOI 10.1038/s41586-026-10652-y. Open access. (Opened and read: abstract, main results, figure legends, methods, discussion, and ethics declarations. Primary source for the Crow, Falcon and Finch architecture, the ten mechanisms and 30 candidates, the LLM-judged tournaments, the 551 papers in 30 minutes and 294-hour estimate, the US$10.76 run cost, the round-one candidates and Y-27632 result, the threefold ABCA1 finding, the 1.89-fold and 1.75-fold ripasudil results, n = 3 and n = 4 wells, the primary-cell validation, the LDH cytotoxicity check, KL001, the Deep Research comparison, and the Edison Scientific competing interest. Every quoted phrase attributed to the paper is verbatim.)
- Nature Portfolio, "Artificial intelligence: AI research assistants that may accelerate scientific discovery," press release, 20 May 2026. (Opened and read. Source for the language models underlying Robin, o4-mini and Claude 3.7, and for the companion Co-Scientist paper published alongside it. No quotation in this report is taken from it.)
- Laura Tran, "An AI-Powered Scientist Proposes a Treatment for Blindness," The Scientist, 9 June 2025. (Opened and read; publication date taken from the page's structured data. Written about the preprint version, a year before the Nature publication. Source for Konrad Kording's comments, the quotation verbatim, and for the FutureHouse team's "move-37" acknowledgment, which I paraphrase.)
- Prior reporting in this publication: Report 022, on why AI-discovered drugs stall after early trials; Report 051, on an earlier autonomous-lab discovery claim; Report 124, on self-driving lab failure modes. (Context only. No claim in this report rests on them.)