> had unusually legible handwriting, but even “easy” early modern paleography like this is still the sort of thing that requires days or weeks of training to get the hang of.
Why would you need weeks of training to use some OCR tool? No comparison to any used alternatives in the article. And only using "unusually legible" isn't that relevant for the… usual cases
> This is basically perfect,
I’ve counted at least 5 errors on the first line, how is this anywhere close to perfection???
Same with translation: first, is this an obscure text that has no existing translation to compare the accuracy to instead of relying on your own poor knowledge? Second, what about existing tools?
> which I hadn’t considered as being relevant to understanding a specific early modern map, but which, on reflection, actually are (the Peter Burke book on the Renaissance sense of the past).
How?
> Does this replace the actual reading required? Not at all.
With seemingly irrelevant books like the previous one, yes, it does, the poor student has a rather limited time budget
To your point about OCR: I think you'll find that the existing OCR tools will not know where to begin with the 18th century Mexican medical text in the second case study. If you can find one that is able to transcribe that lettering, please do let me know because it would be incredibly useful.
Speaking entirely for myself here, a pretty significant part of what professional historians do is to take a ton of photos of hard-to-read archival documents, then slowly puzzle them out after the fact - not by using any OCR tool (because none of them that I'm aware of are good enough to deal with difficult paleography) but the old fashioned way, by printing them out, finding individual letters or words that are readable, and then going from there. It's tedious work and it requires at least a few days of training to get the hang of.
If anyone wants to get a sense of what this paleography actually looks like, this is something I wrote about back in 2013 when I was in grad school - https://resobscura.blogspot.com/2013/07/why-does-s-look-like...
For those looking for a specific example of an intermediate-difficulty level manuscript in English, that post shows a manuscript of the John Donne poem "A Triple Fool" which gives a sense of a typical 17th century paleography challenge that GPT-4o is able to transcribe (and which, as far as I know, OCR tools can't handle - though please correct me if I'm wrong). The "Sea surgeon" manuscript below it is what I would consider advanced-intermediate and is around the point where GPT-4o, and probably most PhD students in history, gets completely lost.
re: basically perfect, the errors I see are entirely typos which don't change the meaning (descritto instead of descritta, and the like). So yes, not perfect, but not anything which would impact a historical researcher. In terms of existing tools for translation, the state of the art that I was aware of before LLMs is Google Translate, and I think anyone who tries both on the same text can see which works better there.
re: "irrelevant books," there's really no way to make an objective statement about what's relevant and what's not until you actually read something rather than an AI summary. For that reason, in my own work, this is very much about augmenting rather than replacing human labor. The main work begins after this sort of LLM-augmented research. It isn't replaced by it in any way.
I do understand that a mere user of e.g. OCR tooling does not perform a systematic evaluation with the available tools, although it would be the scientific way to decide for one. For a researcher, however, the lack of knowledge about the tooling ecosystem seems concerning.
> Granted, Monte had unusually legible handwriting, but even “easy” early modern paleography like this is still the sort of thing that requires days or weeks of training to get the hang of.
He isn't talking about weeks of training to learn to use OCR software, he means weeks of training to learn to read that handwriting without any assistance from software at all.
Most of the library is untranslated Latin. I have a book that was recently professionally translated but it has not yet been published. I’d like to benchmark LLMs against this work by having experts rate preference for human translation vs LLM, at a paragraph level.
I’m also interested in a workflow that can enable much more rapid LLM transcriptions and translations — whereby experts might only need to evaluate randomized pages to create a known error rate that can be improved over time. This can be contrasted to a perfect critical edition.
And, on this topic, just yesterday I tried and failed to find English translations of key works by Gustav Fechner, an early German psychologist. This isn’t obscure—he invented the median and created the field of “empirical aesthetics.” A quick translation of some of his work with Claude immediately revealed concept I was looking for. Luckily, I had a German around to validate the translation…
LLMs will have a huge impact on humanities scholarship; we need methods and evals.
The tendency to reaffirm popular beliefs would make current LLMs almost useless for actual historical work, which often involves sifting fact from fiction.
Handwriting recognition, a classic neural network application, and surfacing information and ideas, however flawed, that one may not have had themselves.
This is really cool. This is AI augmenting human capabilities.
Then again, maybe not; OCR is one of the most worked on problems, so the quality of parsing characters into text maybe shouldn't be as surprising.
Off topic: it's wild to me that in 2025 sites like substack don't apply `prefers-color-scheme` logic to all their blogs.
I’m not a historian. I don’t speak old spanish. I am not a domain expert at all. I can’t do what the author of this post can do: expertly review the work of an LLM in his field.
My expertise is in software testing, and I can report that LLMs sometimes have reasonable testing ideas— but that doesn’t mean they are safe and effective when used for that purpose by an amateur.
Despite what the author writes, I cannot use an LLM to get good information about history.
I get a ton of value out of LLMs as a programmer partly because I have 20+ years of programming experience, so it's trivial for me to spot when they are doing "good" work as opposed to making dumb mistakes.
I can't credibly evaluate their higher level output in other disciplines at all.
> There are, again, a couple errors here: it should be “explicación phisica” [physical explanation] not “poetic explanation” in the first line, for instance.
The image seems to say "phicica" (with a "c"), but that's not Spanish. "ph" is not even a thing in Spanish. "Physical" is "física", at least today, IDK about the 1700's. So, if you try to make sense of it in such a way that you assume a nonsense word is you misreading rather than the writer "miswriting", I can see why it assumes it might say "poética", even though that makes less sense semantically.
I love this line and the “flattening of human complexity into numbers” quote above it. It sums up perfectly how I feel about the whole LLM to AGI hype/debate (even though he’s talking about consciousness).
Everyone who develops a model has to jump through the benchmark hoop which we all use to measure progress but we don’t even have anything approaching a rigorous definition of intelligence. Researchers are chasing benchmarks but it doesn’t feel like we’re getting any closer to true intelligence, just flattening its expression into next token prediction (aka everything is a vector).
https://zwischenzugs.com/2023/12/27/what-i-learned-using-pri...
That's an excellent way to put it. It's the default mode of an LLM. You can ask an LLM for biases, and get them, of course.
E.g., the models won't report that unicorns are real because the majority of the internet doesn't report that unicorns are real. Of course, there may be issues (like ghosts?) where the majority of the internet isn't accurate?
But the gist of its argument just seems to be that they don't know fine details of history, and make the same generalized assumptions that humans would make with only a cursory knowledge of a particular topic. This seems unavoidable for a model that compresses a broad swath of human knowledge down to a couple hundred gigabytes.
Using AI as a research tool instead of a fact database is of course a whole different thing.
E.g. I have this recollection of a quote, slightly pithy, from around the 19 hundreds about hobby clubs controlling social life, maybe from Mark twain, maybe not.
I just cannot come up with the prompt that gets me the answer, instead I just get hallucination after hallucination, just confirming whatever I put in, like a student who didn't study for the test and is just going along with what the professor is asking at the oral exam.
I know this is possible, but the further away I get from my core domains, the harder it is for me to use these tools in a way that doesn’t feel like too much blind faith (even if it works!)
I have always felt that LLMs would fall apart beyond the summarization. Maybe they would be able to regurgitate someone else's analysis. The author seems to think there's some level of intelligent creativity at play
I'm hopeful that the author is right. That truly creative thinking may be beyond the abilities of LLMs and be decades away.
I think the author doesn't consider the implications of broad use of LLM societally. Will people be willing to fund human historian grad students when they can get a LLM for a fraction of the price? Will prospective historians have gained the training necessary if they've used an LLM through all of school?
I believe the education system could figure it out over time. I'm more worried that LLMs like this will be used as further justification to defund or halt humanities research. Who needs a history department when I can get 80% for the cost of a few chatGPT queries?
Good historians? Ehhhhhhh.
The problem is one of trust, and it's very difficult to trust the output of LLMs to be correct/true vs "truthy" without extensive verification that may be either as laborious as doing the original research or that may be difficult or impossible without knowledge and understanding of the internals and sources that may not be available.
A hobby of mine is editing Wikipedia articles about Australian motorsport (yes, I have an odd hobby, sue me).
The vehicles in the premier domestic auto racing category in Australia, the Supercars Championship, are unique to the category. Like NASCAR, they're built on a dedicated space frame chassis with body panels that look like either a Mustang or a Camaro draped over the top.
I'd seen occasional claims on forums that when the organising body was deciding on the design of the current generation of cars, they considered using the "Group GT3" rules that are used for a bunch of racing series around the world (including the German DTM championship, the GT World Challenge events raced across Europe, Asia, and Australia, and the IMSA GTD and GTD Pro categories). If true, it might be an interesting side note to the article about the Supercars Championship.
So I asked Copilot (the paid model) to find articles in motor sport media about this (there are a number of professional online publications that cover the series extensively). It confidently claimed that yes, indeed, there was some interest in using GT3 cars in the Supercars championship, and pointed me to three articles making this case.
The first was an article featuring quotes from the promoter of the DTM series saying what a good idea it was to have a common car across different national series. So the first article was relevant, but didn't actually show that anyone involved in the administration of the Supercars Championship was interested in the idea.
The second and third references were articles about drivers and teams whose core business is the Supercars championship also running cars in the local GT3 championship (while not explicitly mentioned in the article, they do this for a large wad of cash from the rich hobbyists who co-drive and fund most GT3 racing). Copilot's interpretation of the articles was just flat-out wrong.
Yes, this was a sample size of one historical query, but its response was very poor.
While I welcome the rise of parallel shadow institutions as civilization grows spiritlessly utilitarian, the future for common sense looks bleak.
Robert Nozick (in Examined Life) asked how we feel if we found out, say, Beethoven seriously composed music based on a secret formula, which is entire mechanical and required no effort for him at all.
Would we still appreciate the music in the same way? If not, does our appreciation really stem from the fact that we feel he has also struggled like we do, and nevertheless produced something incredible.
I remember as a very small child watching figure skaters on TV and thinking "that's no big deal". And before I started programming: "it's just logic, all very straightforward". But that was before I first entered an ice rink or centre-d a div
Maybe we don't really appreciate something unless we appreciate it is hard in a visceral way.
A historian works with (and may even seek out in musty rooms) primary and secondary sources to produce novel research and interpretation.
An AI is at best limited to ~reading sources that human historians/archivists/librarians have already identified and digitized.
Certainly value to be had here wrt to finding needles in and making sense of already-digitized historical records, but that's more like a research assistant.