Skip to main content
🚀 Taking AI from prototype to production? Find the architecture, GPU, security and governance gaps before they become incidents. Get a Production AI Readiness Assessment
GPT-NL title slide, Developing a Dutch LLM from scratch, at the MLOps Community Amsterdam AI Strategy and Sovereignty Special
AI

MLOps Community Amsterdam: AI Sovereignty and GPT-NL

The MLOps Community Amsterdam sovereignty special at UvA: a policy panel on Europe's compute gap and TNO's talk on training GPT-NL, a Dutch LLM, from scratch.

LB
Luca Berton
¡ 10 min read

On 29 January 2026 the MLOps Community Amsterdam Chapter filled a lecture hall at the Universiteit van Amsterdam for its “AI Strategy & Sovereignty Special”. The evening had two halves. One was a policy panel on Europe’s position in AI. The other was a technical talk from TNO on GPT-NL, a Dutch large language model trained from scratch. I build GPU platforms for clients who have to keep data and models within their own borders, so I wanted to hear both halves.

Luca Berton in the Universiteit van Amsterdam lecture hall with the Welcome to the MLOps Community Amsterdam Chapter slide and the panel on stage behind him

The opening slide: “Welcome to the MLOps Community Amsterdam Chapter”, part of 75k members around the world.

The programme

The agenda slide set out the evening:

  • 18:30: First panel, “Policy, Sovereignty & Power”
  • 19:00: GPT-NL, Martino Mensio and Thanasis Trantas (TNO)
  • 19:30: Break
  • 19:45: Second panel, “Builders, Capital & Execution”
  • 20:15: Networking and drinks

Agenda slide for the MLOps Community Amsterdam AI Strategy and Sovereignty Special above the first panel, seen from the back of a full lecture hall

A full lecture hall at the Universiteit van Amsterdam, with the agenda on screen and the first panel already seated.

Panel 1: Policy, Sovereignty & Power

The first panel was Stan van Baarsen (AI Policy @ AI Plan), Andrew Harrison (Agents & GenAI @ ABN AMRO) and Cees Snoek (Professor of AI @ UvA), moderated by Bauke Brenninkmeijer (Orq / MLOps Community). The names and roles here are as they appeared on the slide.

Panel 1 Q&A slide naming Stan van Baarsen, Andrew Harrison, Cees Snoek and moderator Bauke Brenninkmeijer, with the three panellists seated below

The Panel 1 Q&A slide, with the panellists on stage and the moderator at the front.

The moderator opened with a “Setting the scene” slide that gave the numbers behind the debate:

  • Europe has 5% of global compute and the US holds 75%. China is gaining, and H200 exports are now possible.
  • Europe wants to build 19 AI factories and up to 5 AI gigafactories (100k GPUs) through the AI Continent Action Plan (‘27–‘28).
  • Mistral relies on US partnerships with Nvidia and Microsoft. Aleph Alpha pivoted away. The slide said both lobbied to weaken the AI Act.
  • The Groningen AI Factory will have 2–3.5k GPUs. xAI’s Colossus has 200k.

Setting the scene slide on European AI compute, AI gigafactories, Mistral and Aleph Alpha, and the Groningen AI Factory, above the first panel

“Setting the scene”: the compute gap between Europe and the US in four bullets.

These are the figures and claims as the slide presented them. They framed the discussion; I haven’t checked them independently. I don’t have a recording of the panel, so I won’t put words in the panellists’ mouths. What I can share is what the room asked. Questions came in anonymously through Slido, and the ones on screen when I took my photos were:

  • “What are the main factors that limits the growth of AI infrastructure in EU?”
  • “What will AI change in the cyber security space? How we can become less dependent on the USA + how can I contribute as a startup”
  • “What are the things that European governments should avoid doing in order to allow European AI initiatives to catch up to American companies”
  • “Popular Foundation models are mainly based on the huge global Anglo-Saxon corpus. How are we going to tackle all those little eu languages + eu regulations”
  • “Complete shift of Infra in EU is not pragmatic in short to mid term. Where do we start small which can be sustained…”

Slido Q&A screen with anonymous audience questions about EU AI infrastructure, dependence on the USA and small European languages, above the panel

The audience’s Slido questions. Infrastructure, dependence on the US and small languages came up again and again.

Taken together, the questions were practical. People weren’t asking whether Europe should build its own AI. They were asking where to start, and how to keep it running. The question about small European languages led straight into the next talk.

GPT-NL: developing a Dutch LLM from scratch

Athanasios Trantas and Martino Mensio presented “GPT-NL: Developing a Dutch LLM from scratch”. The agenda listed Athanasios as Thanasis. Both are Scientist Integrators in Intelligent Cyber-Physical Systems at TNO and core developers of GPT-NL. Their outline covered data collection and curation, foundation model pre-training, supervised fine-tuning and evaluation.

GPT-NL title slide, Developing a Dutch LLM from scratch, by Athanasios Trantas and Martino Mensio at the MLOps Community Amsterdam AI Strategy and Sovereignty Special, with both speakers at the lectern

Athanasios Trantas and Martino Mensio (TNO) opening the GPT-NL talk.

The core partners on the slide were TNO, SURF and the Nederlands Forensisch Instituut (NFI, part of the Ministry of Justice and Security). The values slide listed sovereignty, trustworthiness, reciprocity and transparency. The speakers were clear that GPT-NL isn’t trying to challenge the big AI providers. The goal is a model that is useful for Dutch society and understands the Dutch language and its idioms. The target users on the slides included financial, legal, insurance and telecom services, government and social welfare, healthcare, education, safety, security and defence, industry and research.

The summary slide called GPT-NL “a responsible large language model built from scratch”:

  • Data: 450 billion text tokens + 150 billion code tokens, made up of opt-in data, data legally accepted for training LLMs, and non-IP-infringing synthetic data.
  • Performance: “comparable to the Llama2 7B model, GPT-3 175B models” for text generation, summarisation and simplification.

GPT-NL slide: a responsible large language model built from scratch, 450 billion text tokens plus 150 billion code tokens, comparable to Llama 2 7B and GPT-3 175B

GPT-NL on one slide: the data, what it’s made of, and the comparison the team uses.

They also compared GPT-NL with GPT-4, Llama 3, Mistral Large 2, DeepSeek R1 and Apertus on six criteria: whether sources are published, web scraping versus permission, copyright (TDM opt-out), anonymisation before training, output safeguards, and whether the model facilitates data subject requests. On the slide, GPT-NL was the only one marked “with permission only” and “opt-in only”, with “extensive” anonymisation (full anonymisation and contextual privacy filters) and output safeguards that restrict undesirable high-risk use cases. Keep in mind that this is TNO’s own comparison of its project against the others.

Comparison table of GPT-4, Llama3, Mistral L2, DeepSeek R1, Apertus and GPT-NL on sources, web scraping, copyright, anonymisation, output safeguards and data subject requests

TNO’s comparison table: how GPT-NL’s data and safeguards differ from other well-known models.

Data: licences first

In the speakers’ words, the data work starts from following the AI Act. Instead of crawling the web and taking everything, the team has rules about which licences are acceptable. Their copyright spectrum ran from CC-0, CC-BY, public domain and agreements, through MIT and Apache 2.0, to CC-BY-SA and “no licence”. At the far end were CC-NC, CC-ND, robots.txt opt-outs and LLM-distilled data. The slide marked where GPT-NL sits on that spectrum.

The pipeline they described:

  1. Extract and unify raw sources (JSON, Markdown, HTML and other formats) into one representation.
  2. Normalise the text, filter on quality, and detect the language, so the language mix the model learns stays balanced.
  3. Remove personal information, except for public figures.

The architecture slide showed the curation stages built on Hugging Face’s datatrove, with Parquet datasets between stages, running on SLURM. The selected data came from open collections such as EU legal and vocabulary sets, public-domain books and literature, open academic data and open code, plus a filtered, GPT-NL-curated web crawl. On top of that, the team collected Dutch and Flemish datasets themselves: municipal council documents, government announcements, parliamentary records, archives, school content and planning-agency reports. Rights holders contribute through a Content Board, described on the GPT-NL Content Board page.

Pre-training on Snellius

The training part was the most MLOps-heavy section of the evening:

  • Distributed training: data, tensor and pipeline parallelism were presented as “a balancing act between memory efficiency & computational efficiency”. The team used Fully Sharded Data Parallel (FSDP), which the slide described as an upgrade of ZeRO-3: it partitions the parameters, gradients and optimiser states but not the activation memory, and “there will always be some overhead” from GPU communication.
  • Codebase: they compared PyTorch FSDP (adapted from OLMo-core) with Hugging Face Transformers plus DeepSpeed for pre-training on Snellius, the Dutch national supercomputer. The model architecture slide cited OLMo-core.
  • Architecture: a Llama 3-style Transformer in two sizes. The 8B has 32 layers, model dimension 4096, FFN dimension 14336, 32 attention heads and 8 key-value heads. The 30B has 60 layers, 6656, 17920, 52 heads and 13 key-value heads. Both use grouped-query attention, RoPE (θ = 500,000), a KV cache and SwiGLU.
  • Schedule: a short linear warm-up, a long constant learning-rate phase, then a linear cool-down over 15–20% of the steps (the “annealing” phase).
  • Tracking: Weights & Biases managed the lifecycle of the model, data and experiments.

The slide I liked most was called “Babysitting”. Time is money, so pre-training has to keep running. That means metrics and alerts for SLURM jobs, team members taking turns to check 3+ times a day, resuming from frequent checkpoints, and managing disk storage limits. Each epoch took 1.5 months: three epochs, from mid-June to December. When the talk took place (Q1 2026), pre-training was finished and the team was preparing a beta launch.

The “Technical lessons learned” slide is worth repeating almost word for word:

  • Fine-tune to your system’s hardware. H100s? Use FlashAttention 3. Fast interconnect? Use sharding, increase batch size, send full-precision shards over the network. Not a lot of nodes? Don’t shard.
  • Scaling LLM pre-training is “not your typical HPC problem”. There are a lot of knobs, but hyperparameter tuning isn’t feasible at large scale.
  • Community effort and knowledge sharing are crucial.

A box on the same slide put a number on it: a 30B model with 450B tokens equals 34 days of training, or 14 tonnes of CO2, “or 40,000 km by car”.

After pre-training

The final section covered why pre-training isn’t enough: a pre-trained model is a raw language generator that doesn’t follow instructions. Modern post-training stacks supervised fine-tuning, preference tuning and reinforcement learning. The team said plainly that they don’t have the resources to do all of it. Their capability-priority table gave the most weight to general instruction following, GPT-NL’s main NLP tasks, RAG and longer context than the base model, with less weight on chat-style interaction and knowledge recall. The slide said they were explicitly not focusing on coding, maths, reasoning, precise instruction following, safety or agentic use.

Where GPT-NL stands now

I checked the official sources while writing this. According to TNO, the project received €13.5 million from the Ministry of Economic Affairs and Climate Policy (via RVO). The model weights are available under a controlled licence, the source code is published as open source, and part of the revenue flows back to the content creators. The original 2023 TNO announcement gives the project’s background. At the time of writing, the GPT-NL website says version 1.0 is planned for autumn 2026 and is currently available to launching customers. Code is on GitHub and models are on Hugging Face.

I didn’t take notes during the second panel, “Builders, Capital & Execution”, so I’ve left it out of this recap.

My take: sovereignty is an infrastructure problem first

The panel slide and the GPT-NL talk made the same point from opposite ends. Europe is short of compute, and the projects that do exist depend on someone running GPUs reliably for months at a time. GPT-NL’s “babysitting” slide is what that looks like in practice: SLURM alerts, checkpoints, disk quotas and people on rota. It’s operations work, not model research.

From the client work I do, sovereign AI rarely means training a foundation model. It means running open-weight models, fine-tuning and RAG on infrastructure you control, under EU jurisdiction, with an audit trail. In practice that’s an on-premises or EU-hosted Kubernetes or OpenShift cluster with GPU nodes, shared fairly between teams, with model serving, monitoring and data lineage on top. The building blocks are mature now: the lessons on multi-tenant GPUs on OpenShift AI and Slurm for GPU clusters cover the two scheduling worlds GPT-NL also had to bridge. What’s usually missing is the operating model, not the hardware.

That’s why I find the licence-first approach of GPT-NL useful even for teams that will never pre-train anything. “Which data are we allowed to use, and can we prove it?” is the same question regulated companies have to answer before RAG goes to production. For the bigger picture, see my posts on digital sovereignty and EU cloud strategy and geopatriation and sovereign tech stacks. It’s also good to see the community side growing up: the MLOps Community joining the CNCF brings these conversations closer to the cloud-native platform teams who will run the infrastructure.

If you’re planning a sovereign AI platform on Kubernetes or OpenShift, with GPUs on premises or with an EU provider, this is the kind of work I help with.

Free 30-min Production AI consultation

Book Now