Skip to main content
🤖 Running agents for a team, not just yourself? Get an independent review of identity, secrets, failover, observability and governance. Assess your agent platform
Sovereign AI Models: Teuken, Apertus and OpenEuroLLM
AI

Sovereign AI Models: Teuken, Apertus and OpenEuroLLM

The sovereign AI models named at Open Source Summit Europe 2026, Teuken, Apertus, OpenEuroLLM, LLM-jp and SEA-LION, and when to use one in production.

LB
Luca Berton
· 4 min read

One slide at Open Source Summit Europe 2026 in Prague named five models in a single box: Teuken (Germany), Apertus (Switzerland), LLM-jp (Japan), OpenEuroLLM (Europe) and SEA-LION (South-East Asia). It came from the Open Source at a Crossroads keynote by Nithya Ruff and Johan Linåker, under the heading “Countries pushing towards sovereign AI”.

Slide: AI, widening open innovation beyond a single project, with Teuken, Apertus, LLM-jp, OpenEuroLLM and SEA-LION

The slide’s argument: as incumbent vendors widen their grip, countries are building models for autonomy, transparency and localisation. And for AI, collaboration has to reach further down the stack than code: model training, tools and infrastructure, data commons, compute, and skills.

I get asked about these models by clients in regulated industries, usually with the question “can we use one instead of a US model?”. Here is what each of them is, based on public sources, and how I would decide.

Teuken (Germany)

Teuken-7B came out of the OpenGPT-X research project. According to Fraunhofer IAIS, it is a 7-billion-parameter model trained from scratch on all 24 official EU languages, with about half of its pretraining data in languages other than English. The project built its own multilingual tokenizer to make training cheaper for European languages.

  • Two variants were published on Hugging Face: a research version and a commercial version under Apache 2.0.
  • Partners included Fraunhofer IAIS and IIS, Forschungszentrum Jülich, TU Dresden, DFKI, IONOS, Aleph Alpha and the broadcaster WDR.
  • The project was funded with about €14 million by the German Federal Ministry for Economic Affairs, and ended in March 2025.

Where it fits: multilingual European use cases where a small, permissively licensed model is enough, such as classification, extraction and retrieval-augmented answers over internal documents.

Apertus (Switzerland)

Apertus, Latin for “open”, comes from the Swiss AI Initiative of EPFL, ETH Zurich and the Swiss National Supercomputing Centre (CSCS). Its selling point is that everything is open: architecture, weights, training data and training recipes. That is a stronger promise than “open weights”.

From public reporting and EPFL’s announcement of Apertus 1.5:

  • Trained on the Alps supercomputer at CSCS, in 8B and 70B parameter versions.
  • Trained on about 15 trillion tokens in more than 1,000 languages, with around 40% non-English data.
  • Licensed for research, education and commercial use. It is available on Hugging Face and through Swisscom.
  • Version 1.5 adds image and audio understanding, better reasoning, and stronger tool use.

Where it fits: organisations that need to audit what a model was trained on. If your compliance team asks “where did the training data come from?”, Apertus is one of the few models with a documented answer.

OpenEuroLLM (Europe)

OpenEuroLLM is the EU-funded effort, through the Digital Europe Programme, to build open models for all official EU languages, with a consortium of around 20 European partners and compute on EU supercomputers. Public reporting earlier this year described an 8B model targeted for summer 2026 and a larger flagship planned for 2028, under Apache 2.0. Check the project site for the current release status before you plan around it.

Where it fits: long-term strategy. If you are building a platform for the public sector in the EU, plan for a future where an EU-funded model family is the default procurement choice.

LLM-jp (Japan) and SEA-LION (South-East Asia)

The slide also included two non-European examples, and they make the point that this is not a European trend alone:

  • LLM-jp is Japan’s academic and industrial effort to build open Japanese-language models.
  • SEA-LION from AI Singapore focuses on South-East Asian languages, which global models often serve poorly.

The pattern is the same: languages and cultures that are under-represented in global training data get their own models, built in the open. I saw the same theme at KubeCon Japan.

Should you use a sovereign model in production?

My honest answer, after helping teams run models on Kubernetes and OpenShift: sovereignty is a property of your whole stack, not of the model alone. A European model on a non-European API you do not control is not sovereign. A capable open-weight model on your own GPUs, with your own data, often is.

How I would decide:

  1. Start from the task. Many enterprise workloads (extraction, routing, summarisation, RAG) work with a 7B–8B model. That is where Teuken, Apertus 8B and the first OpenEuroLLM models compete.
  2. Check the licence and the data story. Apache 2.0 and documented training data (Apertus) beat “open weights, unknown data” for audits.
  3. Benchmark in your languages. Test on your own documents in your languages. Multilingual training shows up most in the smaller EU languages.
  4. Own the serving layer. Run models with open inference stacks such as vLLM or llm-d on Kubernetes, so you can swap models without rebuilding the platform.
  5. Keep a frontier model for the hard 10%. A pragmatic architecture routes most traffic to a sovereign model and only escalates edge cases, under a policy you control.

As the State of Open Source in Europe report put it in Prague, participation is what turns open source into sovereignty. The same applies to models: using them, reporting issues and contributing evaluation data in your language is how these projects improve.

More from Open Source Summit Europe 2026

Free 30-min Production AI consultation

Book Now