Assuming us Europeans finally get our act together, I think it is better for our long-term future (and the ethical problems) if we manage to get a baseline of training input and data ourselves, from scratch, with everything being ethically sourced.
Oh and, while we're at it, the EU has 24 official languages plus a host of minority languages. Most LLMs focus on the English, German, French and Chinese languages, but everything else is... left behind at best. An European model with actual funding and proper data sources might be able to significantly reduce that.
[1] https://www.taiwannews.com.tw/news/6245677
[2] https://www.theguardian.com/technology/2024/apr/16/techscape...
1. Free of controversy like unlicensed training materials
2. Free of exploitative rlfh loops by people in low-wages countries
3. The leasons learned (and published) from going through the entire training process on "European" hardware: "AI factories" (the term for Slurm HPC/HTC systems with lots of heavy GPU nodes, heavily subsidized by our government [0])
1 and 2 are strong counter-LLM arguments at the moment, and hold back some groups of potential users. Another is energy/water use, so going for maximum green energy would be a nice boon as well. 3 is something I consider to be highly useful for our European identity and "way of the ninja" (for you Naruto fans out there).
Its purpose is not to become some kind of OpenAI or global foundation offering services/tools on that scale.
There is a lot of critisism on this project, not invalid, but mostly based in lack of understanding of what the goals are of the organization as well as the people building the thing.
The people building it, are well aware of how it will be less capable than other LLMs on a general reasoning aspect, not only due to having actually purchased _all_ licensed data that has been used as inputs. Not being a multi billion dollar corporation, this means having very little data and should be an obvious signal to observers that it has not the goal to outdo other models.
In my opinion (personal) its a project that has a learning and demonstration value that is not 'look how well our model performs against others', but still offers value.
1. Huge tax incentives, let the companies get grossly wealthy while paying minimal taxes. Minimum 10 years with clauses protecting "retribution" taxes there after.
2. Tax incentives for the founders/shareholders, just like above.
3. Drop worker protections to a minimum, make it easy to fire people. You only want serious/dedicated employees anyway.
Within 2-3 years there will be at least a trillion dollars looking to get in.
Don't worry though if reading that made you mad. Its absolutely not going to happen. I can think of few things more antithetical to the European ethos than smart skilled people working 80-100hrs weeks with almost no vacation to gas their founders net worth by tens, hundreds, of billions.
I love it! So this is our answer to America and China denying foreigners access to their frontier models.. a massive 13,5M€ founding to develop souvereign european ai, trained exclusively on legally obtained documents and highest moral standards as defined in EU AI Act.
[0]: https://www.quotenet.nl/zakelijk/a71588202/techondernemers-m...
Countries should want control over _where_ the compute is happening rather than _what code_ is running.
What's wrong with a country hosting a Kimi, Qwen or GPT-Oss on their hardware for their government work purpose?
Unlike the US, Europe has no California-level VCs. I don't expect hundreds of billions of Euros to be poured into long-shot projects.
Unlike China, Europe has neither cohesive public investment at the global level nor the drive to grow. Long-term investments have a lot of words, a lot of regulations, a lot of proxy goals, but there is neither a lot of money nor urgency. It was captured by this post: https://x.com/piotrsankowski/status/2065795919623438546
So yeah, both in economy and warfare, Europe dooms itself to be in the hands of the US, China, or a mix of both.
I guess we’re going for GPT2 level capability?
An ecosystem is the tribal knowledge, revolving door of talent, known processes etc.
If the end goal is to make a half assed Dutch speaking model, I think it won’t cut it. I don’t see anyone using it over Gemma 4b that runs on my laptop.
An ecosystem is more durable and has desirable second order effects.
Also, when training models, you create talent that then could go to other countries (brain drain). Restricting that brain drain without imposing authoritarian restrictions on the movements of people seems hard, so it seems hard to keep talent as a competitive advantage. If instead the competitive advantage is datacenters with chips, power capacity building, fast path to building datacenters, I think they are easier to retain while preserving the rights of everyone involved.
> This public investment underlines the importance of an independent, trustworthy and future‑proof Dutch language model.
It does, but not in the way you think it does.
This looks like a good step in that direction.
Why don't they work together on it? Companies like Airbus have already been able to do that with aircraft.
And they're going to train an LLM with all kinds of extra difficulties compared to OpenAI for just 13.5M?
The very first Llama was 16M for one training.
#define(HARMFUL)
[edit] Downvoters please tell me what the problem is with specifying this?
This is not even funny. If you want a competitive AI industry, you need to invest much more heavily in infrastructure first, building models second.