← The Signal Report Work with me

Report 182 · AI in the Lab

What the AI X-ray scientist actually did

An "AI X-ray scientist" ran an experiment at SLAC, and the coverage says self-driving labs are next. The paper is better than the headline: it shows a general-purpose language model doing one hard, boring job on a real synchrotron. It also says, in plain words, what it did not show.

I have spent a lot of hours in front of spectrometers, and the honest truth about instrument time is that much of it is not science. It is alignment: nudging the sample, scanning, squinting at a peak, nudging again. So when a paper says an AI agent can take over alignment on a synchrotron beamline, I am interested for practical reasons. I am also the reader most likely to notice when "aligned a crystal" turns into "runs experiments on its own."

The paper

The work is "An agentic artificially intelligent X-ray scientist" by Zhantao Chen, Arun Bansil, Mingda Li and colleagues, published online on 1 July 2026 in Nature Machine Intelligence (volume 8, pages 1075 to 1086). It is open access, the code and data are on Zenodo, and it is unusually candid about its own limits. In the authors' words, the study does not present "new learning mechanisms or agentic AI architectures." It takes existing language models, gives them tools through the model context protocol, and asks whether they can run one concrete task at a real facility.

The task: determine the orientation matrix of a single crystal on a six-circle diffractometer. In plain terms, the instrument needs to know exactly how the crystal is sitting before any real measurement can start. You find one known Bragg reflection, center it, find a second one, center it, and from those two the software works out the crystal's orientation. The paper calls this "an essential first step in any type of single-crystal scattering experiment." That is accurate. It is a first step.

What the agent really did

Most of the work happened in a virtual beamline, a Python simulator the team built to mimic a real six-circle instrument at the Stanford Synchrotron Radiation Lightsource. There, two models (Claude Sonnet 4 and Gemini 2.5 Flash) each got ten independent runs with hidden, randomly generated motor offsets and deliberately wrong lattice parameters. According to the paper, both "successfully aligned the sample accurately in the majority of cases," usually within 5 degrees, and estimated the lattice constant c within 0.01 ångström in "more than half of the trials." Performance "degraded under more adversarial conditions, particularly when reasoning depth was restricted."

Then came the real beamline, BL17-2 at SSRL, using Claude Opus 4 on a crystal of Co3Sn2S2. The agent picked the (0, 0, 6) reflection first, found it, noticed the sample sat about 1.22 degrees off in one angle (the authors suggest the copper holder's uneven surface), corrected for it, and carried that correction into its search for the second reflection, (1, 0, 10). It ignored a strong diffraction ring on the detector and stayed on the Bragg peak. A separate demonstration on a silicon crystal also succeeded.

That is genuinely useful. Finding peaks on a crystal you have never seen, with unknown offsets, is the part of beamtime that eats hours. A general model doing it from a written procedure, without being retrained, is a real result.

Where the coverage drifted

Northeastern's news story, published 6 August, is a fair read in places. It says the team intentionally used an off-the-shelf model, which is true and is the most interesting part. But four details shifted on the way from paper to press.

"Without a handler." The story says the agent reasons and adapts "all without a handler micromanaging every step." On the real beamline, every command the model proposed was typed into the control terminal by a human, for facility safety. The authors stress that the human was a "passive safety intermediary" who ran the commands "without modification," and I take them at their word. Still, a person at the keyboard between the model and the motors is a different setup from autonomy, and a reader of the news story would not know it was there.

"Learned from the experience." The story says the agent spotted the motor glitch and "learned from the experience." What the paper shows is that it reused an offset correction within the same session. The authors are explicit that this behavior was "observed during a limited number of real-beamline trials and should not be interpreted as a statistically established capability," and that real configurations "were not systematically repeated under identical conditions." One good session is a proof of concept, which is exactly what the paper calls it.

"Utmost precision." The story reports Bansil saying the agent pinned down the crystal's orientation "with the utmost precision." The paper is more careful. On the lattice constant, it says the accuracy "does not yet match the precision typically required in crystallography experiments." It also flags that its own alignment-error metric "may not be the most ideal" for diffraction peaks.

"Broad guidance, then it does its own thing." The guidance was not broad. The Methods describe a long, detailed prompt with step-by-step workflow instructions and tool documentation, refined over many simulated runs where the team watched for failure modes like "unreasonable scan ranges." Behaviors the model kept skipping were "explicitly reiterated" in a second, short prompt. The model chose its own reflections and scan ranges, and the authors are right that this differs from a fixed script. But a lot of human expertise went into that prompt, and that expertise is part of the result.

Why the distinction matters

None of this is a knock on the work. The paper reports its limits openly, which is more than many "AI did science" papers manage. The issue is what a reader takes away. "An LLM can follow a careful written alignment procedure on a real beamline, with a human relaying commands, in a proof-of-concept session" is the finding. "AI scientists can now run experiments on their own" is not.

The gap matters for anyone deciding whether to build this into a facility. The hard engineering questions are the ones the paper names as open: removing the human relay safely, repeating the result across many crystals and sessions, writing prompts for each new instrument (the authors suggest "meta-agents" to help), and making command parsing robust. The current version extracts commands with keyword matching, which the authors acknowledge "can be susceptible to ambiguity or formatting errors." On equipment with motor limits and radiation interlocks, that sentence is the whole ballgame.

How to read the next "AI scientist" headline

What was the task? Alignment, analysis, hypothesis or discovery are very different claims. This one was alignment.

Who touched the hardware? If a human relayed or approved commands, that is supervised operation, even when the human changed nothing.

How many real runs? Simulated benchmarks plus one or two live sessions is a demonstration, not a reliability figure.

How long was the prompt? When the paper describes detailed guidance refined over many failures, the procedure is partly human expertise in written form.

By those questions, this paper comes out well: modest task, honest caveats, open code. The headline just needed to say the same thing the paper did.

Sources

  1. Z. Chen, A. N. Petsch, A. J. Israelski, R. Plumley, L. Shen, C. Wang, C. Peng, Y. Ni, A. Bansil, S. Chowdhury and M. Li, "An agentic artificially intelligent X-ray scientist," Nature Machine Intelligence 8, 1075–1086 (2026), published online 1 July 2026, DOI 10.1038/s42256-026-01261-5. Open access. Data: Zenodo 10.5281/zenodo.20017861; code: Zenodo 10.5281/zenodo.20017991. (Primary source, main text and Methods read in full. Source of every description of the task, the virtual beamline, the models and settings, the ten-run benchmark and its quoted results, the SSRL BL17-2 deployment, the human relay step and its quoted description, the 1.22° offset, the quoted caveats on repetition, precision, the alignment metric and keyword parsing, and the prompt-refinement process. Exact per-model success counts appear only in Fig. 2, which I did not transcribe, so none are given here. Author list as given in the article metadata.)
  2. K. Poltorak, "This AI scientist can run an X-ray experiment. Self-driving labs may come next," Northeastern Global News, 6 August 2026. (Coverage, read in full. Source of the quoted "without a handler micromanaging every step," "learned from the experience" and "with the utmost precision" wording, the latter attributed in the story to Bansil, and of the off-the-shelf-model point.)
  3. MGHPCC, "This AI scientist can run an X-ray experiment," 7 August 2026. (Coverage, read. A short pointer to the Northeastern story; context only.)
  4. Prior reporting in this publication: Report 176, what 37,000 AI agents proposed; Report 124, when the lab robot sleepwalks. (Context only. No claim in this report rests on them.)

Disclosure, plainly: Claude, the model family used in this study, is one of the tools I use in my own work, including research and drafting for this publication. I am a customer of Anthropic's products and have no other relationship with the company, Google, Northeastern, SLAC or any of the authors; none of them asked for, saw or reviewed this report. Nothing here is sponsored and no link earns a commission; here's the full policy.

Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports