back
117 comments
Take a binary array of length N, where N is in the hundreds to thousands range. Choose 2% of the bits to set to 1. Now you have a "sparse array".

Now, you want to use this sparse array to represent a note in a song. So you need every note to consistently map to a distinct* sparse array.

However, you also want to be able distinguish a note as being in one song or another. The representation should tell you not only that this is note A but note A in song X.

How might you do that? Well some portion of the ON bits could be held consistent for every A note and some could be used to represent specific contexts.

Stable and variable bits of you will.

Now if you look at two representations of the note A from two songs you'll see they're different. How different are they? Well you could just count the bits they have in common or not, or you can treat them as vectors. (Lines in high dimensional space) Then you can calculate the angle between those two lines. As that angle increases its easier to distinguish the two lines. They won't ever get to full "right angles" between them because of the shared stable bits, but they can be more or less orthogonal.

That's what's happening here. The brain is encoding notes in a way that it can both recognize A, but also recall it in different contexts.

*But not perfectly consistent, we use sparse representations because the brain is noisy and it's more energy efficient. Pretty close is good enough in the brain and you can encode a lot of values in 1000 choose 20 options.

So we are just walking Lucene indexes?
This really would have been harder for me to understand had I not taken linear and abstract algebra courses a few years ago. That area of maths reused common words like "rotation" but with more generalized definitions, which made it was jarring and confusing to hear and take in at the time. When someone said the word "rotate" my mind as if by reflex was already trying visualize a 3d or 2d rotation even when it made no sense for the problem at hand. Being an English speaker my whole life I thought I understood what a rotation was or could be but I didn't.

Same goes for what's being alleged here: Is there even a way to visualize this that makes mathematical sense? What will be the corollaries to this discovery simply as a result of what the mathematics of rotations will dictate?

Same goes for the ordinary English word "Eigenvector".
And yet the main image on the article illustrates a 45 degree rotation along an axis.

From what I understand, you are saying this rotation is non-intuitive. Could you elaborate more or share some relevant links?

Something to keep in mind though is that in a high-dimensional space, approximate orthogonality of independent vectors is almost guaranteed.
Sure, but the neural activity is actually low-dimensional (see Extended Fig 5e). By day 4, the first two principal components of the neural activity explains 75% of the variance in response. ~3-4 dimensions is not particularly high dimensional.
Do you mean to say that the neurons in the brain are operating in a higher-dimensional space than 3?
No? If the samples are randomly chosen then you'd expect the cosign similarity to be low, but there's no such assumption here, in fact it's the exact opposite.
Can you say a bit more on what that means in this context?
Curious. I cannot understand it clearly. Lets take for example "my wife and my mother-in-law" illusion[1]. It is known for it's property that one cannot see both women at once. If we assume that it has something to do with such a coding in neurons, would it mean that those women are orthogonal, or it would mean that they refuse to go orthogonal?

[1] https://brainycounty.com/young-or-old-woman

Sorry, I'm pretty tired, but I fail to see the relation to this article, how does that example apply?

I thought that was more of a case of a human's facial recognition being a special function, and we're not able to process two or more people's faces at the same time. Like, see the details in them, recognize that it's their face.

You're either looking at one person, or the other, but if you try to look at both of them at the same time, they become "blurry", unrecognizable, even though you remember all the other information about them both.

But that's not related to memory integrity and new emotions/sensations?

Really? I have no trouble seeing both at the same time. Nothing special about it, the angles of their respective faces are different enough that it doesn't feel like there's any interference at all.
Hmmm.. I tried to visualize them both at the same time.. it took some effort, but quickly "oscillating" between the two ended up settling (without a jittery oscillating feeling) on seeing both at the same time. Maybe my brain was playing meta tricks on me though?
Wow. That blew my tiny little mind.

I figured out how to change it at will eventually, if you close your eyes then open them and look at the bottom of the picture first it’s an old woman. Do the reverse and it’s a young woman. Eventually you can do that without the eye closing step but never would I say I could see both at once.

Just rapidly switch.

Very interesting!

I spent 10 minutes staring at that picture and saw only the wife. The mother-in-law never appeared.

This happens to me often.

Wish they would outline the two variants

I only see the young woman before I became disinterested in making the other one happen because why

There was another recent article on applications of geometry to analyse neural mechanisms to encode context. It also mentioned a rotation/coiling geometry:

https://www.simonsfoundation.org/2021/04/07/geometrical-thin...

> The work could help reconcile two sides of an ongoing debate about whether short-term memories are maintained through constant, persistent representations or through dynamic neural codes that change over time. Instead of coming down on one side or the other, “our results show that basically they were both right,” Buschman said, with stable neurons achieving the former and switching neurons the latter. The combination of processes is useful because “it actually helps with preventing interference and doing this orthogonal rotation.”

This sounds like the early conservation of momentum / conservation of energy debates. (Not that they used those words back then.)

I don't remember where I came across this (was probably some pop neuroscience blog or maybe radiolab), but there was some theory about how memories seem subject to degredaton when you recall them a lot, and less so when you don't.

I guess that would sort of be like the opposite of DRAM - cells maintain state when undisturbed, but the "refresh" operation is lossy.

I would expect memories to change more the more they are recalled, just like I would expect a story to change the more times it’s told.
That sounds like the kind of thing they talk about on Hidden Brain (NPR). I think I found it:

https://www.npr.org/transcripts/788422090

Quote (although it’s missing context if the full show):

> Yeah, I think it's really interesting. I think it's really interesting to think about why we do these things, why we misrecollect our past, how those kinds of reconstruction errors occur. And I think about it in my own personal life - I share my memories with my partner. And many of us who have partners, we have these sort of collaborative ways in which we recollect. But those collaborations often result in my incorporating information into my memories that were suggested by this individual, but I never experienced. And so I might have this vivid recollection of something that only my partner experienced because we've shared that information so often. And so that's how we can distort memories in the laboratory. We can just get individuals to try and reconstruct events over and over and over again. And with each reconstructive process, they become more and more confident that that event has occurred.

I'm under the anecdotal and subjective impression that I can do a "brain dump" describing a recently-experienced physical event. But it's a one-shot exercise. Close to read-once recall. The archived magnetic 9-track tape that when read becomes a take-up reel of backing and a pile of rust. The memories feel like they're degrading as recalled, like beach sand eroding under foot, and becoming "synthetic", made up. The dump is extremely sparse and patchy. Like a limits-of-perception vision experiment: "I have moderate confidence that I saw a flash towards upper left". Not "I went through the door and down the hall" but "low-confidence of a push with right shoulder, medium-confidence passing a paper curled out from the wall at waist height, and ... that's all I've got". But what shape curl? Where in the hall? You've whatever detail was available around the moment you recalled it, because moments later extra information recalled start tasting different, speculative fill-in-the-blanks untrustworthy.
> I guess that would sort of be like the opposite of DRAM - cells maintain state when undisturbed, but the "refresh" operation is lossy.

Or like any analog data medium ever :)

it's the theory of re-consolidation

here are some references

https://pubmed.ncbi.nlm.nih.gov/?term=memory+reconsolidation...

Perhaps the Crick and Mitchison theory about why we dream: https://en.wikipedia.org/wiki/Reverse_learning

(AFAIK it's totally wrong, but I really like it anyway. I hope there is another specie in the universe that use it.)

How fascinating, I've experienced this myself to a large degree. I have a few songs that very vividly remind me of certain periods or points of my life. When I play them, I always feel like I'm scratching up the vinyl surface of the memory, and I lose a little bit each time. Rather disappointing :(
This maps wonderfully onto SVD, Neural networks and embeddings.

Word embeddings frequently encode particular traits in different 'regions' of a 256(ish) dimensional space. AFAIK, It is also why we think of element wise addition (merging) in neural networks as an efficient and relatively loss-less computation. The aggregation after attention step used in Transformers (GPT-3) fundamentally relies on this being true.

Although from my reading, there is an inherent assumption of sparsity in such situations. So, is it reasonable to assume that human neurons are also relatively sparse in how information is stored ?

The Nature version is paywalled https://www.nature.com/articles/s41593-021-00821-9

but I found the preprint of the paper on biorxiv.org: https://www.biorxiv.org/content/10.1101/641159v1.full

The abstracts are a bit different so I'm not sure how close the preprint is to the published version.
Never mind the bioxiv link; the Nature article is no longer paywalled, see https://www.nature.com/articles/s41593-021-00821-9.pdf
> They had the animals passively listen to sequences of four chords over and over again, in what Buschman dubbed “the worst concert ever.”

Hahahaha!

I read the abstract and don't really get it. How is this different from saying that a group of neurons A is responsible for memory storage and a group of neurons B is responsible for sensory processing, and A != B? I think I'm misunderstanding this "rotation" concept.
It's a good question. It looks like they actually specifically check for this and show that it's not two separate groups of neurons. Instead a subset of the neural population changes their representation of the input as it moves from sensory to memory, so it's more like a single group of neurons that represents current sensory and past memory information in two orthogonal directions.
> And yet those memories can’t be allowed to intrude on our perception of the present, or to be randomly rewritten by new experiences.

Assumes facts not in evidence? I feel it’s incredibly common for memories to intrude on perception of present, and be rewritten by new experiences.

This makes much more sense than having secret memory cells in neurons.
This is basically just linear algebra.
Squeak squeak
Simple explanation if anyone needs it.

The problem is suppose you have 4 neurons that need to understand a memory and present experience at the same time to make a decision (for example that you see a hot stove and memory that hot stoves hurt). The incoming neurons from memory and experience each have 4 neuron connections, which fire at some rate. Lets represent this as the firing rate of the neurons per second in a 4d vector:

experience: <1.0, 0, 0, 0> (1 pulse per sec on axis 0) memory: <1.0, 0, 0, 0> (1 pulse per sec on axis 0)

If you "add" these together at the downstream neurons, you won't be able to tell which was a memory and which was sensation. A simplified explanation of how neurons work is by combining voltages from their incoming neurons. Example:

downstream sees: <2.0, 0, 0, 0>

upstream could be a memory with <1.0, ...> and experience <1.0, ...>, or memory <2.0, ...> experience <0.0, ....>, or memory <0.5, ....> and experience <1.5, ....>. There are many possible vectors that could "add" to produce the downstream effect, so it makes it harder for those neurons to "learn" the pattern.

As a math equivalence, if I ask "what two numbers sum to 10", there are many solutions (its impossible to disentangle the original numbers).

To make it easier to learn these patterns, what if we used only separate elements of the incoming vectors to represent this information (so the elements of memory and experience could be seperated)?

So some intermediate neurons can transform the representation. We can constructor orthogonal vectors (since the vectors above are sparse):

experience: <1.0, 0, 0, 0> memory: <1.0, 0, 0, 0> => <0, 1.0, 0, 0> experience + memory: <1.0, 1.0, 0, 0>

The "memory" must undergo a "rotation" which moves data into an "unused" portion that won't conflict with the experience neuron firing pattern.

Now downstream neurons can use the data from each (its effectively merging memory and experience without confusing the signal). There is only a single memory and a single experience that combined will give the firing pattern, so the pattern can be learned.

Due to the way linear algebra works, its possible to do this with more complex numbers along arbitrary axes in an n-dimensional space (instead of doing it with a single axis/neuron and all others being zero).

For a physical corollary, imagine two images super-imposed on each other. If they are very distinct, you might be able to infer what the two source images were, but if they are similar it would be difficult. Now imagine a "lenticular" image that clearly displays two images by printing them at orthogonal angles on the medium. You can easily determine what content belongs to which image, but only having a single "print" to store the data (this isn't a perfect anology, but it illustrates the idea):

https://images.app.goo.gl/3dCH7Txigh1adTd66

It’s this blockchain?
Can someone liberate the article from behind the paywall for me?
In mice.
Abstract is mostly readable to a technically person:

https://www.nature.com/articles/s41593-021-00821-9

looks similar to "Near-optimal rotation of colour space by zebrafish cones in vivo"

https://www.biorxiv.org/content/10.1101/2020.10.26.356089v1

"Our findings reveal that the specific spectral tunings of the four cone types near optimally rotate the encoding of natural daylight in a principal component analysis (PCA)-like manner to yield one primary achromatic axis, two colour-opponent axes as well as a secondary UV-achromatic axis for prey capture."

Articles on Quanta magazine have clickbait titles.