back
113 comments
Everyone is excited about LLM abilities to help with language learning, while completely ignoring the fact that for most people LLMs will make the learning unneeded. There will be less experts in the field, and therefore we will lose the part of language and foreign literature understanding not captured by statmodels. Which is a huge part (subtle contexts in poetry, etc)
You aren't wrong, but this has been a dilemma with every new technology. The camera had that effect, modern metalworking had that effect, even tractors had that effect.

It's definitely a problem we should be talking about, but we can't go back in time or remain frozen, the genie never goes back in the bottle. We have to move forward towards the future while salvaging the parts of the past we want to bring with us.

As someone who loves writing in 3 languages (French and German being my native languages, English the third), playing with Claude (and to a lesser degree gpt4) has actually made me play and investigate the nuances of languages so much more. It does a great job of course doing stylistic transformations, which on their own are always stilted, but are great inspiration. But it also does a phenomenal job at explaining the nuances between the language, say when I want to explain why a certain German phrasing “feels” different to me.

Certainly seeing the amount of people never learning the language of the country they emigrated to, this is a problem we already have (and in my situation, never going to country where I don’t speak the language).

I think humans are going to continue nerding with language just as much as they ever have, I do really think it’s an innate drive, and llms are a mind blowing tool to do just so.

> Which is a huge part (subtle contexts in poetry, etc)

Certainly, but from what I remember of my GCSE English Literature back at the turn of the millennium, my fellow students and I didn't understand most of that subtly even when it was famous poets in our native language.

Shakespeare may be unsurprising in this regard given the age (why eye of newt and leg of toad? Some say common names of herbs, others that it's just some amusingly vulgar items), but we were also just as oblivious to the lived experience of being gassed in the trenches as per Dulce et Decorum Est or a cavalry charge as per The Charge of the Light Brigade in English as we would have been if this had been a second language.

I'm not sure about "unneeded". Important motivations why people learn a foreign language are because they want to speak with native speakers in that language without an intermediate (because they move to that country, or because they have a partner speaking that language), or because they find that language interesting/beautiful, or because they want to read or listen to original sources. Machine translation doesn't remove any of those motivations.
> while completely ignoring the fact that for most people LLMs will make the learning unneeded

Good luck using an LLM to talk to strangers in a bar in a foreign language.

most Latin text have been translated once and more then 100 years ago while most ancient Arabic text have never been translated. this is an old problem.

I see AI as savior here esp with reconstructing old languages we only have small amount of text saved

This Dan Brown/William Gibson crossover sucks my soul right out of the petabyte SSD I bought in a dark alley.
I've used ChatGPT 3.5 (not 4, too expensive) to translate most of the Latin writings of Jerome, Ambrose, and Ambrosiaster (from Migne's Patrologia Latina) - the translations have been put in this repo in the public domain:

https://github.com/HistoricalChristianFaith/Writings-Databas...

Some takeaways:

- ChatGPT did excellent with about 3 sentences max at a time. Exceeding 3 sentences would cause it to often truncate the response (e.g. translating 3ish of 5 sentences, or hallucinating more).

- ChatGPT would originally return the translation, sometimes randomly prefixed with a variant of "The translation is" and sometimes wrapped in quotes, othertimes not. Using the function interface to ChatGPT eliminated this problem.

- When it comes to quotations from Bible verses, ChatGPT sometimes "embellished" (not sure what else to call it). E.g. if part of Ephesians 2:7 is quoted in Latin, in the English ChatGPT would sometimes insert Ephesians 2:7-8 in full.

I don't really understand what the value in posting these kinds of takeaways about using GPT3.5 here is. GPT4 is significantly better, and improved models are coming. There's just not a lot of point to benchmarking 3.5 when likely every issue you've pointed out is solved by 4.
Could also try Claude-2, as the OP did.
This is really fun. Google Translate is incompetent at Latin, and I'm informed that so far ChatGPT still makes errors of grammar and word choice when generating Latin.

This experiment helps show we can use GPT-4/Claude to parse and summarize Latin, but doesn't yet show that we can rely on them to the level of a human expert.

I'm confident we'll get there pretty soon - and then will be able to rely on LLMs to generate Comprehensible Input and thereby greatly accelerate language learning.

You can include in the prompt a requirement to highlight sections the LLM was not sure about/needs to be verified.
I think that mostly depends on getting more high quality Latin into the training set, but I'm guessing the new amount of that being generated/discovered is relatively small. Then again, new techniques for training models could prove me wrong.
You get really good results if you prompt it with: “You’re an expert in Latin translation”.
"Many people equate the word "daemon" with the word "demon", implying some kind of satanic connection between UNIX and the underworld. This is an egregious misunderstanding. "Daemon" is actually a much older form of "demon"; daemons have no particular bias towards good or evil, but rather serve to help define a person's character or personality. The ancient Greeks' concept of a "personal daemon" was similar to the modern concept of a "guardian angel"—eudaemonia is the state of being helped or protected by a kindly spirit. As a rule, UNIX systems seem to be infested with both daemons and demons."

that naming convention might turn out to be more prescient than people thought. Can't wait until my Catholic school education pays off and I chant at my computer in Latin

Linux actually got the "demons" right: those manifest as usually hidden activity in someone's brain, and in special circumstances can take control over the entire system (e.g. with a deadlock if it's a neutral demon, or by other means if it's malicious). Those Greek daemons, in contrast, never possess or control anyone: they may inspire, but only if the subject is consciously seeking such inspiration ("the gates must be opened from within").
Latin demonology manuals were about GPT-T and Calude in the first place.
"It's clear that GPT-4 and Claude are skilled translators" on what basis? What makes a predictor LLM better at translating Latin than a system trained specifically for translation?

I'm sure they can do a decent job but it's weird to me that someone would leap to GPT-style tech despite its known tendency to hallucinate/make stuff up instead of translation-oriented tools like DeepL or Google Translate (I say this as someone who despises both of those tools due to their quality issues)

I can't imagine there are vast swaths of Latin in GPT's training set.

Well, that's just it - I use Google Translate all the time to translate historical texts, and for whatever reason, GPT-4 and Claude both work better. Since I deal with texts that feature archaic orthography like the long s (ſ) the main advantage over Google Translate is that LLMs can make educated guesses about what a word should be. But even in terms of pure translation ability — assuming all orthographic issues have been corrected — the LLMs do a better job in the languages I've tested and which I can read (early modern Portuguese, Spanish and French, plus Latin).

The post by David Bell which I linked to gets into this for French - I agree with him that ChatGPT (I guess he was using GPT 3.5) has a tendency to "overtranslate." But it is super impressive as a translator overall IMO: https://davidabell.substack.com/p/playing-around-with-machin...

>on what basis? What makes a predictor LLM better at translating Latin than a system trained specifically for translation?

They just are. Sure it sounds a bit strange if you've never thought about it but they are.

>I'm sure they can do a decent job but it's weird to me that someone would leap to GPT-style tech despite its known tendency to hallucinate/make stuff up instead of translation-oriented tools like DeepL or Google Translate

1. They don't just potentially do a decent job. For a couple dozen languages, GPT-4 is by far the best translator you can get your hands on. Google, Deepl are not as good.

2. Tasks like summarization and translation have very low hallucination rates. Not something to be particularly worried about with languages that have sufficient presence in training.

>I can't imagine there are vast swaths of Latin in GPT's training set.

Doesn't matter. There is incredible generalization for predict the next token models as far as proficiency is concerned. a model trained on 500b tokens on English and 50b tokens of french will not speak french like a model trained on only 50b tokens of french but much much better.

https://arxiv.org/abs/2108.13349

It also doesn't need to see translation pairs for every language in its corpus to learn how to translate that language pair(but this is the case for traditional models too)

I haven't attempted latin translations, but anything from my native language to english and back has been 100% perfect, miles better than anything google translate can do
For some important context, The "Attention is all you need" paper that established the transformer architecture that most LLMs use, is a paper about explicitly machine translation.

It the idea of using transformers for non-translation tasks was only briefly explored at the end of the paper. So it really shouldn't be surprising that LLMs are still good at translating.

Yes, the hallucinations are less than ideal, but the extra freedom is part of what makes their translation abilities so good when they do get it right. And it's not look google translate is completely free of "hallucination" type issues. It's well known that dedicated machine translation models will assume (aka hallucinate) genders when going from non-gendered to gendered languages.

I don't think they're arguing that an LLM is better at translation that an actual translator, just that they are pretty good at it. DeepL and Google Translate definitely also make things up though, so I don't think that's a good comparison...
It's a great question. But note that Google Translate is also trained on "predict the missing token": https://blog.research.google/2022/05/24-new-languages-google... / https://arxiv.org/abs/2205.03983 (search the blog post around “Surprisingly, this simple procedure produces high quality zero-shot translations.”)

This was in May 2022, as part of Google Translate adding support for several low-resource languages (including Sanskrit). I was already very surprised that simply training on predicting tokens does translation so well — then a few months later ChatGPT came out, trained (roughly) the same way and doing a lot of things besides translation.

> What makes a predictor LLM better at translating Latin than a system trained specifically for translation?

Contextual awareness that is baked into the models. Large Language Models are at their core transformation engines. For the operation of transformative text there must be awareness of context. This alone makes LLMs great candidates for translation tasks.

I think it's probably great at Latin and Greek for reasons that should be obvious (plenty of public domain raw material, vast reams of scholarship dating back centuries). It's less good with some other languages, eg some Japanese companies have decided to train their own models due to dissatisfaction with ChatGPT's shortcomings.
Slightly unrelated, since each model is trained and tunes for specific task(s), but the original transformer architecture and paper was built with translation in mind. The original performance tests were language translation benchmarks.
LLM based translates use/add contextual information.

It just choose better word when the original is ambiguous.

Hallucinating in translation task is quite low (much lower than creative, fact finding or information retrieval task)

Translation-specialised models like Google Translate don't actually understand what they're translating. But models like GPT do. This fact is intuitive and easy for anyone to test.
I tried using BingGPT to translate simple Chinese text from screenshots. The results were complete hallucinations, different each time for the same screenshot.

I wouldn't trust these translations at all.

That’s a completely different test. You’re using the vision multimodal ability to decipher Chinese script, essentially adding an OCR step to the process, and it’s not good at OCR of Chinese script.

Try feeding it actual Chinese characters. From what I understand, it’s somewhat competent.

Image input in Bing basically can't handle non English text. Has nothing to do with its Chinese translation ability, which is great.
I envisioned future hacking will be like whispering magic poems and spells aka. prompts to AI systems. I know about prompt injection, however this would raise things to a new level :)
Coming soon, layoffs in medieval history departments.
Gpt4 does an ok job translating texts that aren't complex, but if you read the original and its translation side by side, you'll see that gpt4 still makes dumb mistakes every few sentences, hallucinates stuff when it runs into cryptic words it's not familiar with, and sometimes omits important passages. Gpt4 is like a very productive, but clueless newbie.
I always figured our AI overlords would kill us in a Terminator kind of way, not in a "No! You must not read from the book!" sort of way.

https://www.youtube.com/watch?v=E0DIsPBczcI

I always wondered why these texts are so difficult to interpret... ... ... why certain symbols, like a crow have ambivalent meanings. In some cultures the crow is evil, while in others, it's benevolent.

GPT4 to the rescue, let's see what'll happen if everyone has the means to summon demons, curse others and the like.

What could possibly go wrong.
I've been working on a language-learning app, and gpt-4 has made things doable that didn't seem to be doable without it. For example, translating to lesser known languages like Cree. Online you can find a dictionary of words, or phrases, but I wasn't able to find a translator for arbitrary statements. It's not listed on google translate, nor yandex, et al. But, it's possible with gpt.

My strategy in automated translation is to translate, and then translate back to English. That way we can be fairly sure the translation is accurate. Of course if this app "has wings" I would open source the translations to allow corrections, and/or hire native translators for languages I don't speak (which is the vast majority of languages).

Last tip is that for me, I was able to make a pretty good automated system for this. What I did was spell out a monstrosity of a system-prompt which ensures a few things. One, it will always give me the most descriptive romanization for non-ASCII languages. Two, it will give me output that is essentially .csv data. Three, if it encounters a made up language like Pirate, it will try (and do a pretty good job). The rest is just parsing my final translation file to find prompt/language pairs which aren't saved yet and piping that to the monstrous prompt which queries for translations.

Then you can run a separate program for validation. See if it passes my game of telephone test.

This way I and others could learn common phrases in languages which are hard to otherwise access. It even works for dead languages/dialects/character-voices, etc. For me, it's pretty amazing actually. Please note that I never said perfect. But it's pretty damn close actually.

All in all this workflow allows what would have been unimaginable even two years ago.

In doing tests today I found it interesting/useful to note how gpt-4 is "thinking" about translating the word "settings". Here is its thoughts. This is not the exact system prompt I'm using for my app.

  system_prompt = "You are an expert in translating lesser-known languages. When translating you will include both the native writing system, and the romanization into the latin alphabet. When you romanize text you always include any accents or pronunciation marks."

  user_prompt = "Translate the following into Cree. [Hello, goodbye, settings]"

  # Output
  In Cree language, your words will be translated into:

  1. Hello - ᑌᔭᔭᑎ, romanized as "Tānsi"
  2. Goodbye - ᐊᔭᙱᐂᒥᑎ, romanized as "Ayāwāw"

  "Settings" is a bit more complex, because it implies technological context that doesn't necessarily have a direct equivalent in Cree. However, a possible option is:

  3. Settings - ᓂᐹᕗᓂᑕᐚᓇᐠ, romanized as "Nipāvunitawānāk", which might refer to "adjustments".
I lost three hours reading through the rest of this wonderful blog.
Sure give the AI barely constrained by alignment prompts that flip off when given the wrong prompt arcane rites to summon demons that will end well...
It makes a few mistakes still, and that makes it "just a tool"? Ok, how about in a few weeks when it stops making mistakes?
Real demons prefer Latin.
Oh no, more demons summoned! AI ruins the world.
Institutional review board time. We've already been warned that a computer merely enumerating the names of God can end existence. Then surely a computer can also summon ancient demons.
From what I've gathered from Catholic exorcists, demons adopt different personas, and shouldn't be trusted about anything they say. The only questions the exorcist asks are those pertaining to the case, all in the interest of breaking the claim of demons and expelling them, in the name of Christ, the stronger man from the parable. As the Lord says, Satan is a liar and a murderer from the start, and when he lies, he speaks his native tongue. What I'm saying, keeping a database of demons makes little sense.
Slightly off topic but this is hilarious. We are already crafting "chants" and "spells" for LLM, i.e. prompt engineering, now we are teaching it demonology too? Some priest from the middle ages would have a heart attack.

Now I know how the AI apocalypse would look like. GPT-42 would summon hordes of demons from the pit of Hell to bring about the end of days. Who need all that pesky nuclear codes when you can call upon Satan?

I remember the moment I first saw ChatGPT. "Finally I can translate all my Latin demonology manuals"