Report 170 · AI in the Lab
What 75% accuracy meant in AI mind-reading
In late 2023 a Japanese team reported that generative AI could rebuild images people only imagined, and the coverage ran with "over 75% accuracy." The number was a two-choice test where a coin scores 50%, and it was graded in the same feature space the algorithm had been tuned to match. An independent reanalysis of the released code, posted in November 2025, says the picture gets much weaker from there.
"AI Can Recreate Images From Human Brain Waves With 'Over 75% Accuracy'" was PetaPixel's headline on 4 December 2023. It was a fair summary of what the press materials said, and almost nobody reading it would have guessed what the 75 was a percentage of.
This report is about that number. It is also about a reanalysis that most of the original coverage never followed up on, because a careful critique two years later does not travel the way a breakthrough does.
What the paper claimed
The study is Koide-Majima, Nishimoto and Majima, "Mental image reconstruction from human brain activity," published in Neural Networks (volume 170, available online November 2023). Participants lay in an fMRI scanner, first viewing images and then imagining them. A decoder translated their brain activity into the internal features of image-recognition networks, and a generator then adjusted a picture until its features matched the decoded ones. The abstract states the result directly:
Quantitative evaluation showed that our framework could identify seen and imagined images highly accurately compared to the chance accuracy (seen: 90.7%, imagery: 75.6%, chance accuracy: 50.0%). In contrast, the previous method could only identify seen images (seen: 64.3%, imagery: 50.4%).
Note the phrase "chance accuracy: 50.0%." The authors put it right there. It is the single most important number in the sentence, and it is the one that fell out on the way to the headline.
What "identification accuracy" is
Pairwise identification works like this. Take one reconstruction. Show it the true target image and one other image picked at random. Ask which of the two the reconstruction is more similar to. Repeat across many pairs and report the fraction answered correctly. A process that answered at random would land at 50%.
So 75.6% does not mean the reconstructions looked 75.6% like what people imagined. It means that, in a two-way choice, the reconstruction sat closer to the right answer about three times in four. That is real information above chance. It is not a fidelity score, and "the previous method" at 50.4% was, by the same yardstick, indistinguishable from a coin.
The second question is the one that matters more: "more similar" according to what? The paper's own preprint answers it. The bioRxiv version describes the identification analysis and then says "The weighted similarity in each reconstruction algorithm was used as the similarity metric." In plain terms, the reconstructions were scored with the same feature comparison the algorithm had spent its iterations minimizing.
Why that is circular
A reanalysis posted to arXiv on 11 November 2025 by Ken Shirakawa, Yoshihiro Nagano, Misato Tanaka, Fan Cheng and Yukiyasu Kamitani (ATR and Kyoto University) names this directly. The algorithm "explicitly drives reconstructions to match those features," they write, so "testing similarity in the same space merely verifies the algorithm's internal consistency rather than its perceptual validity."
They then demonstrated it. Using the authors' released code, they rebuilt images from the decoded semantic (CLIP) features alone, dropping the visual features that carry most of the picture's actual structure. The resulting images "differed substantially from the targets," yet pairwise identification scored in CLIP space stayed high, around 75%. Scored instead by plain pixel correlation, which the algorithm never optimized, accuracy fell to near chance.
The spectroscopy version of this mistake is one I have watched people make. You fit a model to a spectrum, then report how well the model matches the spectrum as proof the model is right. The residual tells you the fit converged. It tells you nothing about whether the assignment is correct, because the fit was built to make that residual small. A test only means something if it can fail.
The pictures were the best ones
The reanalysis lists five concerns. The circular metric is one. Here are the others, as the authors state them.
Selective presentation. The reconstructions in the paper's main figures came from a single participant, subject 2. Reconstructions of the same targets from the other participants were in the supplementary material and "appeared noticeably lower in quality." When the reanalysis team generated reconstructions for images the paper did not show, even subject 2's were "qualitatively poorer than those reported."
Run-to-run randomness. The algorithm has more than one source of randomness, and repeated runs on identical brain data gave visibly different images. The team writes that "despite our considerable effort, we were unable to reproduce results that consistently matched the quality of those showcased in the original paper and press release." They raise the possibility, which they frame as a possibility, that the showcased images were the most appealing of several attempts. They also note the code was released four months after publication, so reviewers could not have checked any of this.
Baselines. In a four-way ablation that switched the two headline additions (the "Bayesian" sampling step and the CLIP features) on and off, and scored the outputs with metrics independent of the algorithm, no condition was consistently best. The full method against the condition equivalent to the older 2019 method showed "minimal" improvement.
The Bayesian step that barely moves. The released algorithm runs 1,000 ordinary optimization steps and then 500 steps of the sampling procedure (stochastic gradient Langevin dynamics) that gives the method its "Bayesian" name. The reanalysis reports the sampling step size peaks around 0.00016 and the temperature is set to 10-6. Comparing images just before and just after those 500 steps, across three participants and 25 samples, the mean absolute pixel difference was 4.89 ± 1.00 on a 0 to 255 scale. In their words, the component is "functionally inert."
Where the coverage drifted further
Once a number is loose, it mutates. A January 2024 piece on the Japan Science and Technology Agency's Science Japan site (translated from Japanese trade press) describes participants being asked to imagine images, then reports "successful image reconstruction with a high average accuracy rate of 90.7%." In the paper, 90.7% is the figure for images people were looking at. The imagined-image figure is 75.6%.
The same piece describes the image features fed to the decoder as "inception scores." In the paper the Inception Score is an evaluation metric for how natural a generated image looks, which the reanalysis points out says nothing about whether it matches the target. Small slips, both in the direction of a bigger result.
Who is critiquing whom
This is not a neutral referee, and you should know that. The method the 2024 paper extended is Shen, Horikawa, Majima and Kamitani (2019), built in Kamitani's lab, and the 2024 study used that lab's published dataset. Kei Majima, the 2024 paper's senior author, is a co-author of the 2019 method. So the reanalysis comes from the group whose work the paper presented itself as improving on, and the dispute is partly about credit.
That cuts both ways and does not settle anything by itself. What makes the reanalysis worth taking seriously is that it is checkable: it uses the original authors' own released code, and its analysis code is public. It is also a preprint. As of this report it has one version on arXiv, it has not been peer reviewed, and I could not find a published response from the original authors. If one appears, this report will be updated in place.
Four questions for any "AI reads your mind" headline
What is chance? If the accuracy comes from a two-way choice, 50% is the floor, not zero. A result reported without its chance level is not a result you can read.
Scored in whose space? If the similarity used to grade the reconstructions is the quantity the algorithm was optimizing, the score measures convergence. Ask for an independent metric, ideally human judges.
How many people, and which ones are pictured? Three participants is a normal fMRI study size, which is fine. Showing one of them in the main figures, without saying why, is not.
One run or the best of several? If the method is stochastic, the honest figure is a typical output, or many outputs. The prettiest of several is a different claim.
Brain decoding is real science, and decoding imagined content at all is a hard, interesting problem. That is exactly why its numbers deserve to be read carefully. A coin flip scores 50% on this test, and that one fact changes how "75%" reads.
Sources
- K. Shirakawa, Y. Nagano, M. Tanaka, F. L. Cheng and Y. Kamitani, "Advancing credibility and transparency in brain-to-image reconstruction research: Reanalysis of Koide-Majima, Nishimoto, and Majima (Neural Networks, 2024)," arXiv:2511.07960v1 [q-bio.NC], 11 November 2025. Preprint, not peer reviewed. (Primary source for the critique, opened and read in full, 30 pp. Source of: the five concerns; the subject 2 selective-presentation finding and the quoted phrases about supplementary and unshown reconstructions; run-to-run variability and the quoted reproduction sentence; the four-month code-release lag; the circularity argument and both quoted sentences; the CLIP-only demonstration (about 75% in CLIP space, near chance by pixel correlation); the four-condition ablation and "minimal" improvement; the 1,000 Adam plus 500 SGLD schedule, step size maximum about 0.00016, T = 10-6, and the 4.89 ± 1.00 pre/post mean absolute error across three participants and 25 samples; "functionally inert"; and the authors' affiliations. Only v1 exists as of 1 October 2026. Works it cites, including Kriegeskorte et al. 2009 on circular analysis, were not separately opened.)
- N. Koide-Majima, S. Nishimoto and K. Majima, "Mental image reconstruction from human brain activity: Neural decoding of mental imagery via deep neural network-based Bayesian estimation," Neural Networks 170:349–363 (2024), DOI 10.1016/j.neunet.2023.11.024, PMID 38016230. (Primary source for the claim. The abstract, including the quoted accuracy sentence, was read via the Europe PMC record, which dates first publication to 9 November 2023 and lists a CC BY license. The publisher's full-text page refused automated access, so the published body text was not read; methods details come from the preprint below.)
- Koide-Majima, Nishimoto and Majima, "Mental image reconstruction from human brain activity," bioRxiv preprint, version 2, 28 March 2023, DOI 10.1101/2023.01.22.525062 (the preprint linked to the journal article in Crossref). (Opened and read. Source of the pairwise identification procedure and the quoted sentence that each algorithm's weighted similarity was used as the similarity metric. Wording in the published version may differ.)
- Pesala Bandara, "AI Can Recreate Images From Human Brain Waves With 'Over 75% Accuracy'," PetaPixel, 4 December 2023. (Opened. Coverage example: the headline, and the "75.6% accuracy rate" versus "50.4%" framing with no chance level.)
- "Mental image reconstruction using generative AI: Decoding brain signals into images: New technology developed by QST," Science Japan (Japan Science and Technology Agency), 15 January 2024, translated with permission from The Science News Ltd. (Opened. Source of the quoted 90.7% sentence attached to imagined images and the "inception scores" description.)
- G. Shen, T. Horikawa, K. Majima and Y. Kamitani, "Deep image reconstruction from human brain activity," PLoS Computational Biology 15:e1006633 (2019), DOI 10.1371/journal.pcbi.1006633. (Authorship and date confirmed from Crossref metadata only; the paper was not reopened for this report. Cited for who built the baseline method.)
Disclosure, plainly: I have no relationship with either research group, QST, NICT, Osaka University, ATR or Kyoto University, and no stake in brain-decoding technology. Nothing here is sponsored and no link earns a commission; here's the full policy.