back
350 comments
I look forward to the day where I'm wearing my headphones in a foreign land and hearing all of the discussions in my own language.

The "universal translator" which was part of Star Trek and a lot of other Sci-Fi I was exposed to as a kid was something I was really fascinated with. My Dad worked as a simultaneous French->English translator and sadly spent long hours away from home and, as a kid, I started trying to build a translator so that it could do his work and he could be home more.

Translation is important work and one that could help a lot of people. It's my hope that we get to the point where these models work entirely on locally carried resources.

I worked on building exactly this earlier this year. I was hanging out in Taiwan for a few months and thought, surely the Babel Fish should exist by now.

I did several experiments recording from all the microphones I could on my iPhone and AirPods while out in the wild. My conclusion: it's impossible right now for that hardware given the microphones we have and what they pick up.

So much of what's spoken is at a combination of (a) high distance (b) low volume (c) background obscuration. Something that was clear as day to my ears would barely register on the mics. While context is of course an issue, the raw audio didn't have enough to even translate.

The one caveat is that there might be low-level (i.e., Apple-only) access to headphone microphones that capture the environment to do noise cancellation. I'm not sure though---I couldn't find them on any API.

For cases where you do have clear audio, existing apps (e.g., Google Translate) are so close to achieving this, but don't let you specify audio outputs with enough fine grained control. By default, it will start screaming out of your phone what you were attempting to silently translate.

The problem is you need a full sentence, plus surrounding sentences to properly translate a lot of things (aka context matters).

So no matter what, conversations in your native speech would have to be delayed before translation.

If I am not wrong, Google Pixel buds offer live translate feature.
Another lesson we can learn from Sci-Fi is very often different species on a planet would have their tribal / local languages and dialects but all spoke a common tongue. I think this is the more humanizing approach, rather than delegate even more of our fleshly processing power to machines.
I’m wearing the Rayban Meta right now and they are already mind blowing, I can already talk to that Meta AI assistant seamlessly. I bet one of the future iteration will have exactly this.
I look forward to the day when that problem is solved by a company that doesn’t mine my data to sell ads.
how am i supposed to talk shit with my friends about other people in public then
Babel Fish
> so that it could do his work and he could be home more.

That's a nice way of saying unemployed.

Can't wait for someone to roll a language tutor out with this tech.

Everyone gets a personal tutor for hours a day.

I would absolutely love a VR game where I just need to work in China or Mexico all day and pick up the language that way.

Seamless Streaming looks really promising! We just had a new employee start a few months back with profound hearing loss and our company had no idea what to do with him from an accessibility standpoint. They threw out solutions like Dragon, not realizing those solutions are not real-time.

He ended up rolling his own solution by standing up Whisper in one of our clusters and writing a basic front end and API to take his laptop’s mic input and chunk it every few seconds to send to the model and get back text in pseudo-realtime. We got him a pretty beefy Alienware so he wouldn’t be tied to the cluster GPUs. I can’t wait to see what he does with these new models!

Impressive work, really excited for this.

I will note though that I feel safer getting an occasional bad word than I do having a translator straight up deceive me.

For example, "what the fuck" in English->Spanish is giving "qué diablos" output. Definitely toning down the meaning there.

If someone says something mean to me, I want to know it.

It’s amazing how far text to speech has come in the past few years, but what I’m wondering is when this tech will finally make it into local TTS engines baked into the OS (eg for screen readers, etc)
My wife was training to be a professional voice actor to do dubbing in several languages when we met.

I told her then that the industry would be disrupted by AI before she retired.

Glad she pivoted. Really impressive results.

If "toxic word hallucinations" isn't a cyberpunk phrase I don't know what is.

(quote from the video presentation in the link)

And just the other day StyleTTS[0].

Just text to speech has gone too far. Audio books would be mainly generated on the fly like this?

I think some RPGs in some 5 years time might have something like this:

- A text file that outlines characters and a lose plot/Story line. Human written.

- 3D Mesh Generation based on character description via Transformers based models. Auto generated.

- Dialogues for each NPC via LLM.

- This TTS engine again based on such models.

Result - almost unlimited replayability. Or even edit text file, have a new world based on a new story line with characters having different personas.

[0]. https://news.ycombinator.com/item?id=38335255

I don't see how realtime voice translation can ever be possible; to properly translate the first half of my sentence, you need to hear the whole sentence first. I don't know how simultaneous translators can translate from verb-at-the-end languages like German, until they know the verb.

It's not just where the verb is; sometimes I say something ambiguous, and my next utterance is supposed to acknowledge and remedy that. But if that ambiguity doesn't exist in the target language, I don't see how a simultaneous translator can convey the ambiguity, without knowing how the next utterance is going to refer to it.

Maybe that's why human simultaneous translators often seem to stumble or backtrack. I've never met someone whose job was simultaneous translation. It must be very difficult.

I'm impressed by this effort to convey non-linguistic elements of speech in translation. It's quite an achievement, and a very ambitious goal.

Aside: I wish I knew how speakers of tonal Chinese dialects express feeling, when tonality is supposed to convey semantics. When I hear chinese speakers, I can "hear" the feeling, but I don't know how they do it - it can't just be down to emphasis. (I learned some mandarin 50 years ago, at school. I learned the tones, but they didn't teach expression; and I was never taught by a native speaker, although there were language-lab tapes.)

I had pretty terrible results when I tried English -> Swahili I'm using the Huggingface M4T V2 spaces, it pretty much doesn't work most of the time and I just get English back with a different voice, Expressive on the other hand only has a few languages it seems.

It would be nice if they could layout what exactly is missing in terms of data to make a language work better, while the actual AI bit is out of reach for most of us maybe we could provide more data.

There is also a 60 sec limit and wonder if this is HuggingFace limitation or Seamless?

I'm thrilled to see the progress made in the last 30 years.

As a student in the mid-90s I worked on a system called Verbmobil at the German Research Center for AI and it did speech-to-speech for English, German and Japanese in very limited domain.

This was done via "classical" NLP: You had to model the domain with concepts, you needed sentence parsers, semantic engines, speech-to-text hand-crafted for 3 languages etc.

As it turns out, this approach is/was a dead-end.

Wow, after trying out the demo, I'm floored by how high quality this is. The translations worked perfectly, the voice cloning was "good enough", and the emotions conveyed in my voice was retained pretty accurately.

I don't think this would fool anyone that I was a real native speaker of the target language, but for casual conversation this would work pretty much perfectly. It basically avoids all of the traditional pitfalls of machine translation, like the unnatural robotic voice that it outputs, the slow translation speed and huge latency for realtime conversation, and the loss of emotion.

Yet again, Hindi (the major language in India) is not even in the samples. India is the largest user base of facebook (and probably 1/3rd of the engineers working there are Indians) but never will facebook put enough effort to contribute back. Only use the DAU from India in investor calls.
We can’t be that far off from almost perfect real-time translation. There is some latency of course to hear and process
>Automatically filters out toxic speech >Watermarking

So it can't be trusted at all then

Does the spanish expressive sample sound muffled for others too? And the french sounds super mechanical. Hopefully, it's more impressive the other way.

Also: "This research demo is not open to residents of, or those accessing the demo from, the States of Illinois or Texas"

Can anyone help demystify the licensing?

Besides the ACCEPTABLE_USE_POLICY, there's a CC BY-NC 4.0 (NonCommercial) license, a 'SEAMLESS_LICENSE' (NonCommercial), but also an MIT license? It would seem these other licenses contradict the MIT license, could somebody help clarify how these all interact in practice?

Try the demo here, you record a video of yourself and it does voice cloning and a comparison:

https://seamless.metademolab.com/expressive/?utm_source=meta...

Like how easy it is to get going but you need to download about 20GB and s2st needs 40GB GPU RAM!

It runs but any audio input (you will need to provide wav not mp3's) I tried (tried 20s/40s/300s) I get just one short sentence returned in target language that seems not related at all to my audio input (i.e. Tous les humains sont créés égaux).

Seems like some default text but it runs on full GPU for 10 minutes. Tons of bug reports in GitHub as well.

Text Translate works but not sure what is the context length of the model. Seems short at first glance (haven't looked into it).

Oh and why is Whisper a dependency? Seems not need if FB has their own model?

Any more info about the watermarking? Only Meta can make the determination?

Edit: I can’t find the weights but if I’m reading the paper right anyone could train their own detector.

This tech from Google seems similar, but doesn't have a fancy demo: https://blog.research.google/2023/12/unsupervised-speech-to-...
How will Meta put these models into practice? I understand why Google and Apple have models for their mobile OS users, but I don't understand where users for Meta speech models come from. Are they planning to show Instagram videos with English narration in French or what?
Every video in this page is a bit out of sync with the audio. Combined with the blandness of feature expressions and the whole mood in general, I kept waiting for the moment when the video would disclosure that everything on it was created by AI.
I've been trying (and mostly failing) at settings up a pipeline to get system audio into whisper and feed that transcription into a seamless m4t text-to-text translation model. It seems like seamless streaming is going to solve most of my issues, and should significantly reduce latency!

My ultimate goal is to have realtime translations of video conferences. I've moved to a new country, and while I'm super privileged that most of my colleagues speak English, we still have a number of "all hands" meetings that I get lost in pretty easily.

I was hoping to find out, that the actor's voice in the demo video was generated, or that he had recorded the video speaking in another language or something.

That would have been the knockout punch.

As a French native speaker, I am surprised by the low quality (frankly ridiculous) voice of the French translation example.

Especially because the head of AI at Meta is a French guy AFAIK (Yann Lecun).

Besides the obvious good news about making it easier for people to communicate with each other across languages, it's also exciting to me that we're trending towards a world where I can tap into all the knowledge that only exists on the non-English web. I'm sure there are vast troves of programming knowledge in the Japanese-only web for example. The Chinese-only and Russian-only web are obvious candidates too but presumably those are harder to access for other reasons.
How far from a real-time Star Trek translator? Whisper is fast enough and light enough, LLMs are getting there, so it’s close isn’t it?
How does this compare to whisper-large-v3 on STT?
The demo is so much fun to use. I can't wait for all these technologies to start integrating into filmmaking / games.
Next step is combining the output with few-sample speech synthesis so the output is in the original speaker's voice!
The near-realtime aspect of this is so promising -- we're getting closer and closer to IRL babelfish!

What I would love to see is an ability to add my own voice (yes, at the risk of deepfakes) so that the model could "speak" in any language and sound more like me, not some random voice actor it was trained on.

Neat. How translatable are tones of voice for intent across languages? Like does a person trying to do a "nerdy" voice(nasally, whiny, etc.) in English translate to the "nerdy" stereotype for a French speaker. Seems to do very good on whispers which made me wonder what could be next.
It really sucks that a company so irresponsible with all your data is one of the leading AI companies now.
I want this as a channel in our discord.

Would allow more interactions of people that don’t speak the same language

How did that page get camera access without my permission?

Edit: by the upvote I guess it wasn't just me?

Currently Steam bans games from using AI-generated assets (for good reason). I wonder if they'll back track on this or carve exceptions because this tech seems really useful for indie devs to add voice work to their otherwise silent games.
LICENSE

Attribution-NonCommercial 4.0 International

https://github.com/facebookresearch/seamless_communication/b...

Can this do speech to text English -> English? Get strange results if I do a translation to the same language would be an interesting alternative to Whisper if it could.
It’s funny, all the humanities types try to push the proliferation of languages, but the engineering types keep trying to reduce the language barrier.