back
167 comments
I've got a scan from a book that I OCR with new releases. Ligatures, critical sigla, Fraktur letterforms, subscripts, superscripts, etc.

Nothing special about this model for overly-detailed work like mine.

It's been a while since I last tested (and discontinued my subscription), but the "pro" models from OpenAI dominate. Not surprising, given the price difference, but it would be nice if an OCR-specific model could perform better. It's worth mentioning that even the highest-end models do a pretty poor job with intricate text like mine.

While I haven't tried OpenAI for OCR, I've put my small scale OCR work through both Claude and Mistral OCR. Claude is absolutely better - even in OCR work I did last week and compared with Mistral OCR 4.0.

Mistral's one advantage is that Anthropic now flags OCR, because they don't allow anything that could be considered "reproduction", even of work for which you own the copyright. So my new workflow is Mistral OCR for the actual OCR, followed by a proofreading pass by Claude (which is allowed). Claude is obviously more expensive, but it caught entirely hallucinated sentences created by Mistral OCR 4.0, so I was glad for the backup check.

I have been quite happy with Mistral OCR for the documents I needed to process (typeset, but old, with questionable scan quality, sometimes elaborate typesetting or, much worse, typewriter-and-handwriting approximations of it). I do not test every new model when they are released, but I did a review shortly after Mistral OCR 3 was released and it was a very good compromise: cheap, fast, and good results without further processing. I found generalist models to be way too much faf to get them to avoid unnecessary modifications to the text and report accurate bounding boxes for figures and tables.

That said, models have sometimes surprising weaknesses and a model could be terrible overall but magically work for one type of document.

I got the opposite experience very recently : tried to OCR a bunch of handwritten emails addresses with chatGPT and I had to make so many corrections that I gave up. Whereas Mistral nailed it on first pass.
> the "pro" models from OpenAI dominate. Not surprising considering the price difference, but it would ne nice if an OCR-specific model could do better.

I haven't been impressed with any of Mistral's models. They obviously realized that they couldn't compete at the frontier so they decided to go for smaller focused models but even those have not been that good.

Yet: how is pricing? Evaluating contents and routing appropriately isn't a new challenge in OCR, one of the oldest fields of applications in ML. Thus, how do the smaller open models perform in tandem with relatively pricy $/pg models & APIs? Your use case is remarkably rare relative to the volume and price sensitivity of enterprise data warehouse ops.
I made a benchmark for handwriting recognition for a project while keeping line breaks and errors (grammar, spelling). Sonnet absolutely dominates it since a good half a year. 5.6 did not change that for me. This should also translate to better ocr.
So is Mistral OCR the best one? Have any other OCR models caught some of what you describe? I've been kind of interested in how "OCR" type models work compared to old school OCR.
Do all the models give you the bounding boxes, block labels as this one (allegedly) do?
Whats the best open OCR at the moment?
At this point I lost all hope for Europe playing any significant role in the AI race. If that’s a good or a bad thing I don’t know, but it seems to me like that’s the reality.
Commoditization is a beautiful thing. It seems Anthropic and OpenAI are really struggling to maintain much of a moat. Mistral might not be leading but it's not trailing by that much either. And of course the Chinese are doing their own thing quite successfully.

The reality is that the US is betting its economy on data centers at great expense and is exposing its economy to great risk.

Also while geographically a lot of the money and processing power is in the US, the US has been relying on immigration to power its universities and especially AI research has roots all over the globe. India, China, Russia, Europe, etc. AI related know how is finding its way back to all these places.

So, I'm not too worried about the long term here. It will be interesting to see if Anthropic and OpenAI survive their IPOs. Seems like a risky financial bet at this point given the apparent lack of a moat. But if it works out, it will result in a lot of that IPO money being invested in data centers abroad. Including in the EU. Because data residency is a thing here and the EU is too big of a market for companies with that kind of valuation to ignore. We also produce a lot of energy infrastructure (e.g. gas and wind turbines). Those data centers will need lots of power.

Much like spaceflight, aerospace or nuclear engineering, you need to retain local talent for national defense purposes. Being 70% as good is still way better than being 100% dependent and heavily leveraged by your opponents.
It's not a race. You don't get anything for winning.
Not being a rat in the rat race is the real win.
And US lost the significant role in chip/pc manufacturing. But if a product becomes commodity or utility (which at least for now it seems is the direction), with little lockin, it's not a big deal.

I hope we (EU) don't waste money trying to train local models (which at least some people in Poland try to do), and tries to build our own chips - AI chips have different architecture than regular processor/GPU, and TSMC doesn't need to be winner in this new race.

And if not this, then smaller labs, harnesses and actual application.

Mistral is (wisely) are changing their strategy. They switched to hosting open models and they are investing in hosting inference within EU.

Fine tuning an open model to European values is significantly cheaper than making your own model.

Even Africa will surpass Europe with its massive data center upstarts breaking ground.
> hope for Europe playing any significant role in the AI race

Ha. It's a large market of the LLMs consumption. So it which will affect the AI race. Just from other perspective than you assumed.

The largest non-US non-China model is Russian which surprised me.
Yeah? And here I've been a happy Transkribus customer for some time now. If there are better models or interfaces out there for analyzing historical handwriting, I'll definitely take a look.
They've been unfortunately hugged by the "little death" of working with the EU institutions, though I wouldn't bet on them not coming back.
Huh? Mistral 7b was pioneering in its day and IMO they have been very on top of releasing niche useful models like moderation, OCR, etc.

I’m glad Mistral is working on useful solutions.

OpenAI/Anthropic is like a retarded little sibling chasing “AGI” and giving up on rich media and other modalities.

OpenAI/Anthropic is the worst of the mainstream AI.

It goes:

1. Gemini

2. Vidu

3. Le Chat (Mistral)

4. DeepAI

5. [insert MiniMax provider]

For anyone interested, I have an ocr pipeline running on rented GPUs, doing around 1000pages for 0.05-01 usd with around 0.8 seconds per page with full bounding boxes support for grounding.

If you’re interested you can find contact to me via this profile.

3.5 usd/1000 pages is just too expensive…

Accuracy is truly what people die for in the OCR game. Price isn't the primary function here.. it's an equation of price, accuracy, speed, and in mayn cases regulation.
Can it produce accessible PDF files that will pass accessibility tests? Someone who can do that will make a killing laundering PDFs for academia: by April 26, every PDF, syllabus, and academic document needs to comply with WCAG 2.1 Level AA, which means structural tagging, alt text, and lots of other checklist items that AI could probably generate.
Is it European-hosted and fully outside of both CLOUD Act and CCP reach?

Because I'm assuming that's why they get to charge more for the right type of customer.

You should put contact details in your profile :)
The VLM's are so good at complex document understanding now. But you just can't trust them not to invisibly censor sensitive clinical/legal docs, even at the maximally permissive settings.

And the deep learning OCR-only models won't censor, but can and do hallucinate. I've yet to see a 'scan with different approaches and reconcile and say you're not sure if they don't agree' system just work for generic complex documents.

1000 Pages / 3.5€ this is expensive as hell. If this is not fastly superior than something like tesseract it is not worth it.
Does anyone know a site that lets you browse examples of input / output pairs?, particularly with layout analysis (bounding boxes of figures, tables, etc).
I think people misunderstand the utility of Mistral's OCR. It's not going to beat SOTA models for extraction on edge-case docs, but it's MUCH cheaper and faster and does an excellent job on simple ones. I've been working on converting PDFs to EPUBs and Mistral has been making steady improvements. On a chapter of Bleak House it was able to extract and tag the header, titles, and references at the bottom every time. The only thing it struggled on was line numbers in the right margin which it correctly tagged as "aside text" 3/5 times, but always separated from the core text each time. The important thing to keep in mind is that there's no prompting needed, just upload the PDF and voila!. There's even a batch mode with a 50% discount.
I won't comment on accuracy, but in internal benchmarks, Mistral OCR is significantly faster than comparable APIs.
How does this compare to Baidu Unlimited OCR. I've been very impressed with Baidu and it's essentially free to run on a decent computer, other than electricity costs.
Given the latest vibe release's new "follow default" model option, we should see a new coding / general Mistral model, mistral 4, soon.
I've been experimenting with using NuExtract this week on locally OCRing bank statements that don't have a predefined document structure. It's way better than Tesseract or a generic vision-enabled model. It runs great on a single RTX 4090 at the modest throughput I need.

Their hosted, API-based service is something like a third of the cost of this model.

Mistral has been very good with handwritten ocr, expecting the new models to get better with that across languages.
Benchmarks? I'm currently using Tesseract (via OcrMyPDF) and would like to compare the difference.
I'm wondering if this model performs better on french (and other european languages) documents than others
How does this perform with non-Latin and non-LTR scripts? Say, Chinese, Arabic, Devangari, Adlam, etc.?
Mistral is bumping the price of this thing every release. I think we're at 2x now?
How does this compare to 4?
The chinese did it better, mistral is alive thanks to regulations.
I didn't quite understand.
Who the hell at Mistral thinks it is a good idea to register CMD + T as a shortcut for switching theme!?
Whoever is paying all that for OCR is being scammed.