NEW YORK — When people wanted to create better artificial intelligence, they looked to the brain for inspiration. Now the tables are turning, and researchers are using AI models to better understand the brain. But how do we judge which AI models might accurately describe the computations the brain is performing?
Together with colleagues at Ben-Gurion University of the Negev and Université du Luxembourg, Nikolaus Kriegeskorte, PhD, a principal investigator at Columbia's Zuckerman Institute, explained emerging methods to evaluate which models best capture what computations the brain actually performs. We sat down with Dr. Kriegeskorte, also the Zuckerman Institute's director of cognitive imaging and a professor of psychology and neuroscience at Columbia, to talk about this research published August 28 as a review in Nature Reviews Neuroscience.
Why have scientists developed all these models of how the brain works?
Our brains have billions of neurons, and neuroscientists want to learn how all these brain cells work together to enable us to see, hear, think, plan and act in the world. This has led researchers to develop many different theories about how our brains enable us to do these things. To test these theories, researchers implement them in computer models and see how these models behave.
Can these models perform human tasks?
Yes, they can perform many tasks about as well as humans can. But we want to find the models that achieve this performance using the same computations as the human brain.
Consider face recognition. Scientists often use natural stimuli, such as photos of actual faces, to test both people and models. The idea is to see how well the models do in the real world compared with real people. But if different models all perform the same when asked to distinguish one face from another, we have no way of deciding between them. Some of the models that solve the problem may be doing so in ways entirely different from how our brains solve the problem.
How are you distinguishing the competing models?
We and others have started designing synthetic visual images that are optimized with AI methods to make the models disagree. In these instances, different models make distinct predictions about human behavior. We can then show these images to people or to animals in experiments and see which model is correct in terms of best matching their responses, or at least closer to being correct. This way we might find out what the brain is really doing.
The gist is that neuroscientists are creating artificial stimuli specifically designed to make the models disagree in their predictions. Such "controversial" stimuli – controversial among the models – put us in a good position to adjudicate between computational theories.
neuroscientists are creating artificial stimuli specifically designed to make the models disagree in their predictions.
What is an example of a controversial stimulus?
Let's say you are performing an experiment where people try to identify an image of a number. You could measure their accuracy and then compare how well the models do at predicting whether people think a given number is a 3 or a 7, for example.
When you show people normal images of numbers, one problem is that all the people in the study are probably going to classify all the digits nearly perfectly, and all the models are going to perform similarly well. So then it’s difficult to tell which model is correct. We want to instead present something where the models disagree. We show ambiguous images where one model says this is a 3 and another model says this is a 7. Then you show these ambiguous stimuli to people and see what judgment they come to. You can see which models align best with human behavior.
Does this research strategy apply to more than visual perception?
It doesn't have to be about looking at images. It could be about any sensory process, or many other kinds of processes the brain performs.
What kinds of theories of brain computation have you tested?
Let’s talk about recognizing different handwritten digits. One way in which we can imagine how our visual system works is that it classifies digits based on their features, such as expecting that a 7 has one sharp angle, a diagonal line and a horizontal line. This is called discriminative recognition. But a contrasting theory claims that the visual system uses generation for perception: We imagine different hypothetical things that could be out there in order to recognize which of them is actually out there. For instance, rather than recognizing a 3 by just its most distinctive features, we have a mental model of 3s. The model more comprehensively captures all the images that we might say are 3s, despite variation among them in terms of their precise curves or the thicknesses of their lines.
How did you test such abstract theories?
We compared different computational models that implement particular versions of these theories. Some of the models were purely discriminative; others incorporated generative computations into the perceptual process.
My collaborator Tal Golan at Ben-Gurion University of the Negev in Israel created synthetic images that these models classified differently. For example, we made an image that to one model looked like a 7 and to another model looked like a 3. Both models are highly confident in their judgment, and virtually certain it’s not the other digit. We made such images for each pair of models and each pair of digits, creating a large set of ambiguous images that were controversial among the models. We then showed these ambiguous images to human subjects and had them tell us what digits they recognized.
What did you find?
We learned that generative computations play an important role in visual recognition. Our visual systems somehow get the benefits of both theories. On the one hand, we can get to a rough guess of what’s in front of us within a small fraction of a second using rapid discriminative computations. But purely discriminative neural models could not explain the human judgments of ambiguous images. Our results suggested that the visual system combines discriminative recognition, which is computationally efficient, with generative computations that enable us to learn more from less sensory data. I now think generative computations are important not just for imagining things, but also for perceiving things.
What happens in a situation where none of the models you test explain human behavior?
When you are pitting two models against each other, you know that they both can't be right, because they're making diverging predictions. But they could both be wrong. So, one model might see an image as a 3, and another might see it as a 7, and to people, it doesn't look like anything. You've found a weakness of both models that tells you we need a new model. In fact, we usually find that all our current models are wrong.
Does it discourage you when all models turn out to be wrong?
There is something bittersweet for a scientist when this happens. It’s challenging when our favorite models fail. But we are not discouraged because some models are clearly less wrong than others. These models point us in promising theoretical directions.
Dr. Kriegeskorte is also a member of the Alan Kanzer Center for Cognition and Reasoning, a new interdisciplinary research initiative at the Zuckerman Institute.

