back

by benbreen·1y ago·view on hn ↗
I agree, I probably should've gone into more detail on the actual case studies and implications. I may write this up as a more academic article at some point so I have space to do that.

To your point about OCR: I think you'll find that the existing OCR tools will not know where to begin with the 18th century Mexican medical text in the second case study. If you can find one that is able to transcribe that lettering, please do let me know because it would be incredibly useful.

Speaking entirely for myself here, a pretty significant part of what professional historians do is to take a ton of photos of hard-to-read archival documents, then slowly puzzle them out after the fact - not by using any OCR tool (because none of them that I'm aware of are good enough to deal with difficult paleography) but the old fashioned way, by printing them out, finding individual letters or words that are readable, and then going from there. It's tedious work and it requires at least a few days of training to get the hang of.

If anyone wants to get a sense of what this paleography actually looks like, this is something I wrote about back in 2013 when I was in grad school - https://resobscura.blogspot.com/2013/07/why-does-s-look-like...

For those looking for a specific example of an intermediate-difficulty level manuscript in English, that post shows a manuscript of the John Donne poem "A Triple Fool" which gives a sense of a typical 17th century paleography challenge that GPT-4o is able to transcribe (and which, as far as I know, OCR tools can't handle - though please correct me if I'm wrong). The "Sea surgeon" manuscript below it is what I would consider advanced-intermediate and is around the point where GPT-4o, and probably most PhD students in history, gets completely lost.

re: basically perfect, the errors I see are entirely typos which don't change the meaning (descritto instead of descritta, and the like). So yes, not perfect, but not anything which would impact a historical researcher. In terms of existing tools for translation, the state of the art that I was aware of before LLMs is Google Translate, and I think anyone who tries both on the same text can see which works better there.

re: "irrelevant books," there's really no way to make an objective statement about what's relevant and what's not until you actually read something rather than an AI summary. For that reason, in my own work, this is very much about augmenting rather than replacing human labor. The main work begins after this sort of LLM-augmented research. It isn't replaced by it in any way.

1 comments
> To your point about OCR: I think you'll find that the existing OCR tools will not know where to begin with the 18th century Mexican medical text in the second case study. If you can find one that is able to transcribe that lettering, please do let me know because it would be incredibly useful.

My point about OCR is you haven't done any comparison and is now making the same mistake of claiming without any evidence. The most basic one from google translate does "know where to begin", it even doesn't make the "physical" mistake, though makes others. It also does know where to begin with the image from your second post, although it seems worse. And it's not the state of the art, and I don't know what that is for spanish either, but again, that wasn't my point. You do not have a care-free option, to be able to understand that "physical" mistake you'd still need to read the source, which means you still need those days/weeks of training

> none of them that I'm aware of are good enough to deal with difficult paleography

And you haven't demonstrated anything re. difficult paleography for the LLMs in your article either!

> entirely typos which don't change the meaning

First, you'd need to actually demonstrate that, and that would require the full accounting which you haven't done (and no, I don't plan to do that either) This could be a typo in a name or a year, which is bound to have some impact on a historical researcher? He'd try searching for a misspelled name and find nothing while there could've been an interesting connection in some other text?

>translation, the state of the art that I was aware of before LLMs is Google Translate, and I think anyone who tries both on the same text can see which works better there.

Yes, do try it, for example, in Deepl, to see that it's not any worse

> no way to make an objective statement about what's relevant and what's not until you actually read something rather than an AI summary

Sure, but presumably you've done that before making the claim of relevance "on reflection"? So how is it relevant to demand this "human labor" of the students?

Such a confusing comment, because when I enter the text from case study #1 into Deepl, it's very clearly much worse than what Claude or GPT4o can come up with (the first few lines from Deepl are: "With his expositions to all the Tables, particularly of the quality of the countries, et of the most notable things, to be found in them. Which Tables, can be, and t are taught to reduce' together" and so on).

Likewise with using Google translate on both case studies #1 and #2 - the results are self-evidently far worse. In both cases there were multiple errors in each line and in case study #2 it was entirely unable to transcribe or translate the title line. If you see this, please email me at bebreen [at] ucsc dot edu to share the better results you are seeing because I genuinely am interested and open to using alternative tools - I just am not seeing what you are seeing, apparently.

In terms of typos not changing the meaning, yes naturally a real human needs to double check absolutely everything if it's being used in research. We agree on that - the point is simply that this significantly speeds up the initial research process, not that it replaces the expertise necessary to, for instance, double check that a name or year is transcribed correctly. A huge amount of historical research is simply about skimming through documents looking for relevent info to zero in on - this is where LLMs can really help.