As neuroscience gains the power to record thousands of neurons at once and compare biological brains to artificial ones, it finds itself rich in tools but poor in consensus — more than thirty competing methods now claim to measure neural similarity, yet the field has not agreed on what similarity truly means. This is not merely a technical dispute; it is a philosophical one, echoing the oldest challenge in comparative science: whether resemblance is a surface pattern or a window into shared mechanism. The kidney researchers of the 1960s did not stop at correlation — they used it to reveal how
Neuroscience Grapples With How to Compare Brains and AI Models
Neural similarity scores matter only insofar as they illuminate how computation actually happens.
Why does it matter whether we can compare brains to each other or to AI models? What changes if we get this right?
It changes everything about how we understand what brains do. Right now we have these powerful recording technologies and these powerful AI systems, but we're essentially speaking different languages about them. If we can't rigorously compare them, we can't ask whether they're solving problems the same way—and that's the question that leads to real insight.
But you have thirty different methods already. Why not just pick one and move forward?
Because they answer different questions, and we don't always know which question we're asking. Some methods tell you whether one system can predict another. Others tell you whether they organize information the same way. Those are completely different things. If you pick the wrong metric, you get the wrong answer.
So the problem is that people are using the wrong tool without realizing it?
Partly. But it's also that many of these tools are mathematically equivalent—they're the same thing dressed up in different language. The field has created the appearance of choice when there's actually much less diversity than it seems. That's confusing in a different way.
What would it look like if the field got this right?
You'd report multiple metrics, each one capturing something real about how the system works. You'd use proper mathematical distances, not just scores. And you'd always ask: what does this number tell me about mechanism? Does it help me understand why the brain does what it does? If the answer is no, the metric doesn't matter, no matter how novel it is.
Is that happening now?
Not yet. The field is still comfortable with leaderboards and single rankings. But the tools exist. It's a question of whether we have the discipline to use them properly.
Il Polso
- Over thirty competing methods for measuring neural similarity have flooded the field, leaving practitioners without a principled way to choose among them.
- Popular tools like RSA and CKA are routinely treated as distinct when they are mathematically equivalent — a hidden redundancy that masks a deeper conceptual confusion.
- Predictive accuracy and geometric similarity are being conflated, yet a model can perfectly forecast brain activity while organizing information in an entirely different way.
- Most similarity scores lack the mathematical structure of true metrics, meaning they produce numbers but not navigable maps — they cannot support clustering, embedding, or coherent comparison across systems.
- The field is gravitating toward leaderboard culture, rewarding novel metrics and high scores over the harder work of mechanistic understanding.
- Researchers are beginning to argue that multiple complementary metrics, treated as proper mathematical distances, offer the only honest path toward genuine insight into neural computation.
As neuroscience gains the power to record thousands of neurons at once and compare biological brains to artificial ones, it finds itself rich in tools but poor in consensus — more than thirty competing methods now claim to measure neural similarity, yet the field has not agreed on what similarity truly means. This is not merely a technical dispute; it is a philosophical one, echoing the oldest challenge in comparative science: whether resemblance is a surface pattern or a window into shared mechanism. The kidney researchers of the 1960s did not stop at correlation — they used it to reveal how the body actually works. Neuroscience now stands at the same threshold, and the choice it makes about how to measure likeness will determine whether it catalogs or truly understands.
Comparison has always been neuroscience's quiet engine. Darwin built evolution with it. Kidney researchers in the 1960s compared medullary thickness across desert and non-desert mammals and, in that simple measurement, found a clue to how countercurrent multiplication actually works. Good comparison does not merely catalog — it builds mechanism.
Neuroscience has long compared brains anatomically, noting that hippocampal size tracks spatial ability or that certain structures appear across species. But a harder question has only recently become technically possible: when you record from hundreds of neurons simultaneously in two different animals, or compare a biological brain to an artificial network, are they actually computing in the same way? Recording technology has matured, AI systems have grown to resemble brains in their distributed architecture, and the question is now urgent. The field has responded by generating more than thirty different similarity metrics — geometric approaches, predictive approaches, and hybrids — none of which agree, and few of which most practitioners have time to fully understand.
Beneath this proliferation, surprising unities hide. Representational similarity analysis and centered kernel alignment, treated as separate tools, are formally equivalent after a single mean-centering step. Many other methods collapse onto a handful of underlying mathematical objects. But recognizing this hidden unity also exposes a more serious confusion: the field has been treating predictive accuracy and geometric similarity as interchangeable. They are not. A neural network can forecast brain activity with high precision while organizing that information in a completely different shape. Conflating the two questions produces answers that feel rigorous but mislead.
There is also the matter of mathematical structure. Most similarity measures are scores — bare numbers. True metrics are symmetric and obey the triangle inequality, which means they can anchor a genuine map: embed brains and networks into a shared space, cluster them, apply standard tools, ask new questions. The difference between a score and a metric is the difference between a ranking and a landscape.
The field's deepest problem may be the simplest to name: no single number will capture how two neural systems relate across all conditions and experiments. The honest path forward is to report multiple complementary metrics, each illuminating a different facet of computation — not as a leaderboard of novelty, but as a set of lenses aimed at mechanism. The tools exist. What remains is the will to use them for understanding rather than for rankings.
Neuroscience has always relied on comparison. Darwin used it to build evolution. Kidney researchers in the 1960s compared the thickness of the renal medulla across desert and non-desert mammals and found a tight correlation with the animals' ability to concentrate urine—a clue that unlocked understanding of how countercurrent multiplication actually works. Comparison, done well, moves beyond mere cataloging. It builds mechanism. It explains why things work the way they do.
Neuroscientists have long compared brains too. We use anatomical atlases to identify matching structures across species. We note that hippocampus size correlates with spatial navigation ability. But until recently, we lacked the tools to do something harder: record from hundreds or thousands of neurons simultaneously and ask whether two brains—or a brain and an artificial network—are computing in the same way. The technology has arrived. Recording devices are cheaper, smaller, more standardized. We can now capture neural activity at a scale that was impossible a decade ago. And we have built artificial intelligence systems that, at least on the surface, resemble biological brains: distributed computation across large populations of simple units.
But we have not solved the fundamental problem. When you record from the same brain region in two different animals, how do you know if their neural responses are truly alike? When you compare a biological brain to an AI model, what does similarity even mean? The field has generated more than thirty different methods to answer these questions, and they do not agree. Some methods, like representational similarity analysis and centered kernel alignment, treat the problem geometrically—asking whether two systems arrange their responses in the same shape. Others favor prediction, measuring similarity by how well one system's activity can forecast the other's. Still others blend both approaches. A practitioner trying to navigate this landscape faces a bewildering choice, and many simply do not have time to understand the mathematical nuances of each approach.
The proliferation is partly an illusion. Dig beneath the surface and surprising patterns emerge. Representational similarity analysis and centered kernel alignment, routinely treated as separate tools, are formally equivalent once you add a mean-centering step to the first. Many other methods collapse onto a few underlying mathematical objects. Understanding this hidden unity helps. But it also reveals a deeper confusion: the field has conflated different kinds of questions. A predictive score tells you whether one system carries the information needed to reconstruct another. A geometric score tells you whether two systems organize information the same way. These are not the same thing. A neural network might predict brain activity perfectly while organizing that information in a completely different way. Conflating prediction with similarity invites error.
There is also the question of what counts as a proper metric. Many similarity measures are merely scores—numbers without mathematical structure. True metrics are symmetric and obey the triangle inequality: if system A is close to system B, and system B is close to system C, then A and C cannot be far apart. This sounds pedantic, but it is the difference between a number and a map. When a measure is a true metric, you can embed brains and networks into a common space, cluster them, and apply standard machine-learning tools. You can navigate coherently. You can ask new questions.
The deepest challenge, though, may be the simplest to state: brains are complex. No single metric will capture similarity across all experiments or between a model and a biological recording. The field should report multiple metrics, each capturing different aspects of neural computation. This requires hard work—digging into mathematical details and assumptions. But the reward is real understanding, not just rankings. The kidney researchers did not care about medullary thickness for its own sake. They cared because it pointed toward mechanism. Neural similarity scores matter only insofar as they illuminate how computation actually happens. The field has grown comfortable with leaderboards and novel metrics that are technically different but scientifically marginal. Both habits venerate the scores themselves over understanding. That needs to change. The tools exist. The question now is whether neuroscience will use them to build genuine insight.
Citazioni salienti
Neural similarity scores matter only insofar as they point us toward computational mechanisms.— Alex Williams, Flatiron Institute