Ben Hayden
@benhayden
Professor of Neurosurgery, Baylor College of Medicine
Finally, we find a similar representation in large language models, although with the key difference that there is a much stronger distance effect in the LLMs, suggesting one way in which brain math is different than LLM math.
Moreover, the same neurons represent numerosity for counting dots and Arabic numbers, but the brain uses entirely unrelated codes (both simplicial), confirming an observation made by Kutter et al. that MTL uses unrelated codes for symbolic and non-symbolic numbers.
We find plentiful evidence for simplicial structure. For example, numbers closer in numerical space are not closer together in the neural space. Just look at this insane manifold!!
The resulting manifolds, even those composed of random vectors, will, generally speaking, have a simplicial geometry (a generalization of a triangle or tetrahedron to higher dimensions). 9 numbers > 8-plex. Dimensionality is a blessing not a curse! www.annualreviews.org/content/jour...
We take on a very simple question - how the brain represents the numbers from 1-9. Standard approaches in neuroscience look for the implementation of a number line, a hypothesized low-dimensional manifold where neural distance corresponds perfectly with numerical distance. doi.org/10.1126/scie...
Two words with similar meaning will evoke similar patterns, and a third word with very different meaning will evoke a different pattern. But the individual neurons involved are doing totally different things. Meaning is conveyed by patterns, not special neurons. www.biorxiv.org/content/10.1...
I believe we are in the middle of a scientific revolution on the topic of arealization. @srheilbronner.bsky.social and @myoo.bsky.social and I define the Arealization Paradigm as the idea that cytoarchitecturally defined brain areas are the best guide to the functional organization of the brain.
If you think NHP is different, your reading of the literature may not be very up to date. There's LOTS of evidecence that Everything is Everywhere in NHP! We have a whole box devoted to this:
Technically, you are right, but just look at this!!!! To a first approximation, Everything IS Everywhere!
Finally, in 3-person conversations we found the brain uses the same subspace rotation principles for binding meaning to different speaking partners, with greater rotation between self and other than between specific others.
If I mention my hand, Im not talking about your hand. But if I mention the moon, it’s probably the same moon you mentioned. Semantic-identity binding varies by word. And we found that subspace rotation angle does so too. Body parts were most speaker-specific; verbs and function words were least.
Technical aside: Cross-speaker semantic tuning is more orthogonalized than within-speaker half-split estimates of tuning, which controls for lot of thing, like imperfect embeddings, differential fit quality for speaking and listening, and, of course, neural variability, which is very high.
Here’s where manifolds save the day. If the brain can build cross-speaker vectorial semantic representations very carefully so that they live in an intermediate space between fully orthogonal and fully collinear, you can generalize and differentiate at the same time. www.cell.com/cell/fulltex...
Not only that, but we find overlapping semantic tuning at the single neuron level. A neuron that responds to hearing the word “dog” will tend to have a stronger response for speaking the word “dog.”
We converted all the spoken words to semantic embeddings using BERT (with speaker tokens) and regressed firing against embeddings to get semantic tuning curves for each neuron. Not surprisingly, neurons have maximal encoding around the time of speaking and a few hundred ms after words heard.
We used fancy microphones to record participants having conversations while recordings hippocampal neurons. We transcribed and diarized every word spoken and heard. Automated methods aren’t accurate enough, so our stellar team did it, painstakingly, by hand, using Praat.
Love this paper. I just want to highlight this sentence from the abstract:
If you had predicted this 20 years ago they would have laughed at you and called you crazy. (And laughed at you again when you predicted the mouse would be the most important organism in behavioral neuroscience)
It uses vague, unsubstantiated concerns that we will "miss out on innovations" and then offers its methods as a monitoring solution to a problem it pretends to identify but does not.
In fact it's a good thing because they should spend their time doing science, not writing it up. This paper sounds the alarm over hallucinations and mistakes, but, despite lots of data, doesn't show that those are on the rise in LLM-assisted papers. That's sensationalism.
Finally, we find direct links between contextualization and next-word prediction. For example, hippocampus neurons encode the three most likely upcoming words, with strength corresponding to their likelihood.
To implement contextualization, you need positional encoding - a map of where words are. We find that too. Both linear encoding (not too surprising) but also, superimposed, a sinusoidal (periodic) encoding, which was proposed in Vaswani et al.
This pattern explains contextualization for many types of contextualizing pairs, including, for example, adjectives and nouns that they modify. But not all! For example, verbs contextualize objects, but we don't see evidence that subjects contextualize their verbs!
Here's is the key figure from the paper, where the inferred APGs look really similar, even though they are generated from entirely unrelated data (and three more examples):