Weizmann AI tool rebuilds what you see from fMRI brain scans

A team at the Weizmann Institute of Science in Rehovot, Israel, led by computer scientist Michal Irani, has built an AI tool that guesses what a person is looking at from a brain scan and re-creates the image. It also works the other way: given an image, it predicts the brain activity that image would produce. The work was presented at the Cognitive Computational Neuroscience conference in New York.\n\nNeuroscientists have tried for years to use functional magnetic resonance imaging (fMRI) to reconstruct what people see. Early attempts were blurry. fMRI uses a giant magnet to track the flow of oxygenated blood, and the areas that light up are thought to be the most active. The signal is coarse: in typical scanners a highlighted voxel covers around three cubic millimeters, and each cubic millimeter holds around 16,000 neurons. Irani's team used newer, higher-resolution datasets, collected by other researchers, in which a voxel covers around one cubic millimeter.\n\nIrani says earlier tools are not good enough. Show a person a banana and they will produce an image of a banana, but one that “wouldn't have the same structure, the same position.” The new system tries to fix that with a “brain decoder” that has two branches: one predicts the structure of an image (where the colors are, for instance) and the other predicts its content, such as a bunch of bananas on a plate. Those predictions feed a diffusion model, the type of AI best known for making images by gradually cleaning up a noisy mess of pixels, which produces the final reconstruction. The decoder was first trained on existing data from eight people, each of whom had been shown around 9,000 images in a high-resolution fMRI scanner.\n\nThat was far less data than the models needed, so the team trained a second model, an encoder that predicts brain activity from an image, and used the two together. Start with a new image, say a leopard. The encoder predicts what a person's scan would look like on seeing it, and the decoder then reconstructs the image from that predicted scan. At first the result may look little like a leopard, but repeated training this way brings dramatic improvements. The loop lets the team train on as many images as they like, including ones never shown to anyone in a scanner. Irani says around 70% of the training data comes from images not originally paired with fMRI scans.\n\nCombining data from several studies also let the team identify brain regions that seem to share functions across all individuals. One region seemed to respond to food images, another to sports. Irani is now working with neuroscientists to see whether the tools can reveal new things about the brain.\n\nThe resulting “universal brain encoder” works on a scan from a new person with minimal calibration. Other tools typically need about 40 hours of fMRI data on a new person; Irani says her decoder needs only one hour. Neuroscientist Tommy Sprague of the University of California, Santa Barbara, says that matters because fMRI time costs something like $600 to $1,000 an hour, and such tools could speed up research.\n\nThe system is not perfect. Over a Zoom call Irani showed a cake reconstructed as a pile of three sandwiches, and a dog in a bathtub reconstructed as a similarly colored goat in a bathtub. Still, in a comparison test the tool was found to be much better than previously described ones. “All in all, really we outperformed the others by a significant margin,” Irani says. She calls “mind reading” a “cute, jazzy name” for what they are doing.\n\nIrani plans to move beyond images to video and audio, and wants to reconstruct what people are thinking about or imagining, and the contents of their dreams. “That's something we don't have yet,” she says. She hopes a tool like this could let completely paralyzed “locked-in” people communicate through brain activity alone, and could help scientists study mysteries such as what PTSD flashbacks look like. Judy Illes, a neuroethicist and professor of neurology at the University of British Columbia who was not involved, calls the work “magnificent” and says its therapeutic use for people with neurologic conditions is “tremendously exciting.”\n\nPrivacy worries follow. Sprague says the results seem very impressive, but a way to surreptitiously extract what someone is thinking about would let “150 years of sci-fi” come true, which he finds worrisome. A willing person lying still in a scanner is hard enough to arrange, but Irani and other scientists are working on similar approaches to decode brain activity from EEG, which uses a cap of electrodes or even headphones. Sprague thinks Irani's approach would probably “work quite well” at predicting images a person is thinking about but not looking at. Marcello Ienca, a neuroscientist and philosopher at the Technical University of Munich, calls the move to EEG a “game changer”: once a device is calibrated to a user's brain, it could be relatively easy for companies to extract additional information, potentially without consent. He can also imagine some courts allowing mental image reconstructions as legal evidence. Irani acknowledges the potential for misuse with EEG but is not concerned for now: “I'm trying to think only of good things.”
Key facts
- Michal Irani and colleagues at the Weizmann Institute built an AI that reconstructs the image a person is viewing from an fMRI scan, and can also predict brain activity from an image.
- A two-branch decoder (structure and content) feeds a diffusion model; a paired encoder lets the team train on images never shown in a scanner, which make up around 70% of the training data, per Irani.
- Irani says the decoder needs one hour of fMRI data from a new person, against about 40 hours for other tools.
- In a comparison test the tool was found to be much better than previously described ones, though it still fails, for example turning a cake into three sandwiches.
- Scientists including Tommy Sprague and Marcello Ienca warn that similar decoding from EEG could expose inner thoughts, potentially without consent.
Why it matters
Reconstructing what a person sees from fMRI is an old goal that earlier tools served poorly: they could produce a banana, but not one with the same structure or position as the original. Irani's approach separates structure from content and feeds both into a diffusion model, and a paired encoder lets the team train on far more images than were ever scanned, around 70% of the training data per Irani. The reported drop in calibration, one hour of fMRI data on a new person instead of about 40, is what Sprague says makes such tools valuable to neuroscientists, since scanner time runs something like $600 to $1,000 an hour.
Who it affects
Neuroscientists are the first audience: Sprague says the cheaper calibration could speed up research, and Irani is working with neuroscientists to see whether the tools reveal new things about the brain. Irani hopes the approach could eventually help locked-in, completely paralyzed people communicate, and Illes sees therapeutic promise for people with neurologic conditions. Everyone else is touched by the privacy question: Ienca notes companies could extract additional information from a calibrated EEG device, potentially without consent, and he can imagine courts accepting mental image reconstructions as evidence.
How to use it
This is research, not a product. For researchers, the practical change is the calibration burden: Irani says the decoder needs one hour of data from a new person where other tools typically need about 40 hours. The team trained on publicly available fMRI datasets from other researchers, including newer high-resolution ones where a voxel covers around one cubic millimeter. Irani plans to extend the work to video and audio and to reconstruct imagined content and dreams, which she says is not yet possible.
How solid is it
The reporting is MIT Technology Review's account of work presented at the Cognitive Computational Neuroscience conference in New York. The performance claims rest largely on Irani: that the tool needs one hour of data, that about 70% of training data was unpaired with scans, and that it outperformed others by a significant margin. The article reports that in a comparison test the tool was much better than previously described ones, without giving a numerical score. Illes, who was not involved in the research, calls it “magnificent.” Irani herself shows failures, such as the cake rendered as sandwiches and the dog rendered as a goat.
Risks and caveats
The tool reconstructs what a person is looking at; it has not reconstructed thoughts, imagery or dreams, which Irani says she does not have yet. The privacy concern is about where the approach could go. Irani and other scientists are working on similar decoding from EEG, collected through a cap of electrodes or even headphones, and Sprague says that as models improve this will become easier, so the field must take ethics more seriously. Ienca says it could be co-opted for ethically and societally problematic commercial uses. Irani acknowledges the potential for misuse with EEG but says she is not concerned for now.
“None of us can afford 40 hours of imaging for a new subject”
— Tommy Sprague, neuroscientist at the University of California, Santa Barbara