June 17, 2026 · 1 min read

The unglamorous part of sovereign AI

National language-model projects usually stall on who owns the training text and who can allow its use, not on compute or architecture.

Written by Matteo Gasser

A national library has the texts. A broadcaster has decades of transcripts. A ministry has the budget and the press release ready. From there, the model is supposed to write itself. It doesn’t.

We have watched this happen enough times to know where the months go. They do not go into picking an architecture. The recipes are published, and anyone can read them. The compute can be rented in an afternoon. What slows everything down is a question that sounds like paperwork, until it stops the whole project: who actually owns this text, and who is allowed to say yes to using it.

That is rarely one person. The library digitised half its holdings under a grant that has expired. The broadcaster owns the words on screen but licensed the music under them separately. Somewhere there is a reuse clause from the nineties that nobody can read without a lawyer, and the lawyer wants to be paid before reading it. Each set of texts that looked free comes with old agreements, and the agreements don’t match.

GPT-NL has had to work through exactly this kind of thing. So has most of Europe’s wish for models that speak in a local voice and answer to local institutions.

So the team trains on whatever the lawyers clear in time, and the result quietly reflects which archives could handle the paperwork fast. The label says sovereign. The text behind it says something narrower.

We keep expecting the bottleneck to be technical. It almost never is.