back

by padolsey·3y ago·view on hn ↗
> Rather than train the AI to recognize whole words, the researchers created a system that decodes words from smaller components called phonemes. These are the sub-units of speech that form spoken words in the same way that letters form written words. “Hello,” for example, contains four phonemes: “HH,” “AH,” “L” and “OW.”

This is nifty. Also I'm oddly comforted by the fact that this system doesn't "read thoughts". It just maps slightly upstream from actual speech to the relevant speech/motor production regions of the brain. So no immediate concern for thought hacking...

Separately, this makes me wonder what such a system would be for deaf people (with signing ability) who have lost their ability to move their arms. I imagine–optimistically–that one could just attach the electrodes to a slightly different area in the motor cortex and then once again train an AI to decode intent to signs (and speech). So basically the same system?

3 comments
I think not all deaf people have the same mapping of words and their precise phoneme to the typical expected muscle movements. If this mapping differed from a regular person this system would not be that useful on the interpreting speech side. I think we've had pretty good gesture recognition for a while, on the other hand. I bet it's possible to decode individual signs right now but sign language also has a different grammar from typical spoken English and a lot of meaning is context based so it might be tricky in that way, more of a translation problem.
Oh yeh definitely. I meant more specifically: might it be possible to capture the electrical signals (much like this current system) from the parts of the motor cortex that create the series of muscle movements that form a 'sign', and then creating a 2d projection/display of these muscle movements and then ... downstream, a gesture recognition solution as u mention (a big downstream challenge).

It sounds like a lot. Was just a thought experiment of how such spinal blocks/paralysis would affect deaf people and how they'd be able to continue to communicate with their deaf partners. Definitely niche but nonetheless interesting, and I think possible using the same general approach as the OP article.

But yeh to then translate gestures to speech is a distinct and incredibly challenging problem on its own as you allude to. Perhaps in the future they can tap into signing/speaking-translators' brains and have AI learn those mappings in a fuzzy way.

> I think not all deaf people have the same mapping of words and their precise phoneme to the typical expected muscle movement

In fluent sign language, there is something analogous to phonemes. In linguistics these days, they're just called phonemes, and considered equivalent to spoken language phonemes. They're a fixed class of shapes and locations. They combine in certain ways that make up morphemes, which then make up words. It does work very similarly, perhaps identically, to spoken language.

The distribution of handshapes and they way they interact resembles spoken language. For example, it's somewhat hard to say "strengths" and people often produce a slurred "strengfs". The way it slurs together is rather predictable. It's very hard to say "klftggt", and so it just doesn't occur in natural language. Same with sign languages and hard-to-sign combinations.

Phonemes have an exact realization, but they also exist relative to each other, the distance and direction between them is important. This is probably part of why an American can fairly easily understand New Zealand English, despite nearly all of the vowels being different. Another analogy: in tonal languages, if there's a low flat tone, then 3 rising tones, then a low flat tone, that final low tone may be quite a bit higher than the first low tone -- but it will be interpreted as a low tone, as it is judged relative to the previous rising tone. Vowel qualities besides tone work the same way. And so do hand gestures.

There is a lot of variation by dialect/region/community in sign languages. More than in a language like English. This makes it more complicated, but it shouldn't be insurmountable. And of course, not all deaf people speak sign languages as their native language. They would struggle just as people who learn any other language later in life do.

>I think not all deaf people have the same mapping of words and their precise phoneme to the typical expected muscle movements. If this mapping differed from a regular person this system would not be that useful on the interpreting speech side.

...unless it could be trained on the individual?

This is nifty. Also I'm oddly comforted by the fact that this system doesn't "read thoughts".

If you're like most people, you "talk to yourself" internally all the time, so this doesn't seem that far from reading your thoughts.

Actually, I think it's been found out recently that thinking that everyone has an "inner voice" is a presumption by those who do. Apparently it's only 25%~ of the population or so that actually have an inner monologue.

It got revitalised again recently resulting in much commotion from people on either side of the fence: shock that someone might never hear their own voice or talk to themselves inside their head, and shock that someone might have this voice talking to themselves, for somebody who has only known silence.

I think there are a couple studies that back it up, but a lot is anecdotal as people describe their side of the fence.

I have an inner voice myself, but I have thoughts and sensations that my inner voice does not voice, and some where it does.

The human brain is so freaking cool.

But as far as I know no motor signals are sent to my mouth when I talk to myself (i.e. internal monologue), so this wouldn't read my thoughts. I'm not sure what you're saying.
They might not be acted upon, but they could still be sent, up to a "gatekeeper" brain segment.
Is it really that common? I certainly don't.
Do You Have an Inner Voice? Not Everyone Does https://science.howstuffworks.com/life/inside-the-mind/human...
That would be interesting. I would think if you are conscious you have to have an inner voice.
Nope. Thoughts, yes, but you don't by any means need to think in words or audio all the time.
The notion of 'thinking in words' reminds me of how LLMs work, but it certainly isn't me.
If we're talking about actual vocalizations then I do occasionally but it's just stuff like "whoops!" or "god damn it!" even if there's no one around.
A ML system which skipped the brain and just read the physical movements and converted them to voice goes a long way (I understand there are camera and glove based apps that can do this, but I'm not sure what the accuracy is like)