Ultimately, you need to know the environment of whales. Sure, you know the next probable token, and even relationships between tokens, but how do they relate to what whales are doing?, in other words what does it mean?.
It reminds me of the great quote by Feynman: "look at the bird"[1]. You can know the name of a bird in many different languages, but that tells you nothing about it. If you want to understand birds, you need to look at it and see what it's doing.
If we want to communicate with another species, we need to understand what they're doing, why they're communicating, and from there we can begin to infer what they're communicating. It could be simply a callsign/identification/name for each individual, it could be communicating you've found a place with food, it could be communicating about predators, who knows. There could be all sorts of interesting variation and nuance to their calls. But ultimately it's necessary to look (in the general sense of having information about what it's doing -- could be through all kinds of sensors!) at the ~~bird~~ whale :)
[1] Feynman tells the as a lesson from his father. (that you can know their names and still know nothing about the birds themselves).
https://www.youtube.com/watch?v=ga_7j72CVlc
Bonus related video!: https://www.youtube.com/watch?v=M1TiXLGqlM4
Could we build a chat GPT for whales to use, with this approach, sure. Does that get us any closer to translating it. NO, because without shared context to start building connections and relationships its highly unlikely that were going to find commonality.
You can do both. Models trained to receive and predict audio tokens alongside text are coming. You could just add such data as part of training.
>Does that get us any closer to translating it. NO, because without shared context to start building connections and relationships its highly unlikely that were going to find commonality.
If you trained it alongside human languages then shared context(if available) will be figured out by the model. They'll coalesce in the same shared space the way they do for human languages. Enough to translate without examples instead of just speak? Perhaps - it's not like there are examples in the training set for every lang to lang combination modern models are capable of translating.
But whales? What we are (probably) proposing is putting two languages side by side with NO context and saying "figure it out".
Here is a great example of people and language and perception: https://news.mit.edu/2023/how-blue-and-green-appeared-langua...
Im fairly confident that unless we give an LLM a LOT of context between data sets to work with that it won't make much progress at all.
1. Many such concepts cannot be directly mapped between different languages especially with distant language pairs.
2. Language Models and Image models even if both are only trained on their respective modality learn structural representations so similar, you can connect them with a simple linear layer. That's it. Entirely different modalities https://arxiv.org/abs/2209.15162
https://arxiv.org/abs/2304.08485
I think you are severely underestimating the extent to which representations can group in neural networks.
They're mammals, they're social animals, they see. You're assuming a level of alienness we don't have the knowledge or understanding to truly ascertain.
while ww2 code breaking contained other human contexts, perhaps like-context observations could allow for certain categories that might reveal linguistic insights of whales. like whale sounds observed consistently near boats, divers, danger, eating, play, etc.
no clue about how LLM/AI systems process meta contextual training data inputs (is this even a valid phrase about AI?! haha) but i hope whale linguists have considered such an approach.
It works for EN/FR, less well for EN/DE... because English has a huge overlap with these two (more French than German) for its word by word translation. With something like Chinese, and enough time it might be able to guess its way in due to overlaps in the training data.
If you took two completely random corpus's from two languages, and jammed them into this model and handed that system to someone who only knew one language, the approach would fall flat on its face.
It’s like “guessing” the encryption key of a one-time pad.
One plank of wood is not enough to determine solely if it was intended to be part of a house or part of a boat
Humans: “We come in peace. We seek to understand and communicate with you.”
No response.
Humans: “We repeat. We come in peace…”
WhaleGPT: “too late”