📊 Full opportunity report: AMÁLIA · The Three Hard Questions. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Portugal’s AMÁLIA, a €5.5M state-funded European Portuguese LLM, is now operational. However, critical questions about its openness, native-language data, and optimization goals are still unresolved, reflecting broader issues in European sovereign-LLM development.
Portugal’s €5.5 million investment in the AMÁLIA large language model has resulted in a functional, publicly accessible model, marking a significant milestone for the country’s AI efforts. However, critical questions about the model’s openness, native-language data, and strategic goals remain unanswered, raising concerns about the broader European sovereign-LLM landscape.
AMÁLIA, developed through a consortium involving approximately 60 researchers from Portugal’s top research institutions, was officially launched in October 2025. It is based on a continuation of the EuroLLM multilingual foundation, with the Portuguese component trained on about 5.8 billion tokens, including data from Portugal’s national web archive. The model currently outperforms previous open models on Portuguese benchmarks and beats Qwen 3-8B on most tests, though it still trails on some specific tasks.
Despite these technical achievements, questions remain about how open the model truly is, the sufficiency of native-language data, and what the model’s primary optimization goals are. These issues are part of a broader structural challenge facing European sovereign-LLM initiatives, which are often evaluated as individual launches rather than as part of a collective strategy. The final version of AMÁLIA is scheduled for release in June 2026, and the project team has indicated ongoing work to address some of these questions.
AMÁLIA
The three hard
questions.
Portugal spent €5.5M to build a European Portuguese LLM. The base version is operational, the benchmarks beat Qwen 3-8B on most pt-PT tasks. So why are the most important questions still unanswered?
Last month, Duarte O.Carmo published the sharpest public analysis of AMÁLIA — Portugal’s state-funded European Portuguese large language model. He prefaces his critique with the necessary diplomatic apparatus before doing what almost nobody else in the European-sovereign-LLM discourse has been willing to do publicly: asking hard questions about whether the work, as released, actually does what it set out to do. This piece is a structural extension of his analysis. The AMÁLIA case study exposes three hard questions every national LLM effort needs to answer publicly — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
Three questions every national LLM effort needs to answer publicly.
Duarte O.Carmo’s framing maps cleanly onto the structural argument. Each question lands specifically in AMÁLIA — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
The three questions form a structural feedback loop. Q3 (optimization target) determines Q2 (data volume needed) which conditions Q1 (openness sufficient for community contribution). The European sovereign-LLM movement collectively benefits from these questions becoming standard methodology disclosure, not exceptional critique.

Portuguese Flash Cards – Learn Portuguese Language Vocabulary Words and Phrases – Basic Language for Beginners – Gift for Travelers, Kids, and Adults by Travelflips
PORTUGUESE FLASH CARDS – Basic Portuguese words and phrases for beginners and travelers
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
107 billion tokens. 5.8 billion clearly pt-PT.
The structurally tractable question with a structurally surprising answer. For a model whose entire stated purpose is European Portuguese prioritization, the native-language share of extended pre-training is 5.5%. The implications cascade into every other question.
AI model openness evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Olmo standard. AMÁLIA’s current state.
Allen Institute for AI’s Olmo project defines what “fully open” operationally requires. Olmo doesn’t lead frontier benchmarks. That’s not the point. The point is to be the structural reference for openness. AMÁLIA’s “fully open source” claim should track to the operational standard.
native language data training datasets
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Four strategic positions. AMÁLIA between two and three.
Approximately €100M+ in publicly disclosed European sovereign-LLM funding across the major initiatives. The structural question every project faces: what is the actual competitive position you’re staking? Four options — none mutually exclusive — but each requiring different commitments.
European sovereign LLM development books
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three standards. For AMÁLIA and the movement.
The structural critique generalizes beyond AMÁLIA. Italy, France, Germany, Switzerland, the OpenEuroLLM consortium, and every subsequent national project benefit from public discourse holding national LLM efforts to operational standards on openness, data accounting, and strategic positioning.
The European sovereign-AI agenda is a serious strategic project that deserves serious public discourse. O.Carmo’s analysis is what serious public discourse looks like. Appropriately diplomatic. Structurally rigorous. Willing to ask the hard questions in public when the public investment justifies it. More of this is needed — across every European sovereign-LLM project, not just AMÁLIA.
Implications for European Sovereign-LLM Strategies
The development of AMÁLIA exemplifies both the progress and the persistent challenges faced by European countries in building independent, native-language AI models. The unresolved questions about openness, native data sufficiency, and strategic goals highlight systemic issues that could influence policy, funding, and collaboration across Europe. These questions matter because they determine whether European models can achieve genuine sovereignty and competitive performance, or whether they remain experimental efforts without clear strategic clarity.
European Sovereign-LLM Development in Perspective
Across Europe, nations like Italy, Germany, France, and Norway are investing in their own language models, often with similar structural approaches—either training from scratch or building on multilingual foundations. The €5.5 million Portuguese project is part of this broader movement, which aims to reduce dependence on US and Chinese models. However, public discourse often focuses on individual model launches rather than the systemic issues of transparency, native-language data use, and strategic objectives that underpin these efforts. The case of AMÁLIA is particularly notable because it involves a significant public investment and a national commitment, making these questions especially urgent.
“The three questions—openness, native data, and objectives—are fundamental to understanding the true potential and limitations of European LLMs.”
— Duarte O.Carmo
Unanswered Questions About AMÁLIA’s Openness and Goals
It is not yet clear how open the AMÁLIA model truly is in terms of access and transparency, nor how much native Portuguese data is deemed sufficient for strategic purposes. Additionally, the specific optimization goals—whether for general performance, domain-specific tasks, or national sovereignty—remain unspecified. The final version scheduled for June 2026 may address some of these gaps, but current details are limited.
Next Steps for AMÁLIA and European Sovereign Models
The project team plans to release the final version of AMÁLIA in June 2026, which may clarify some of the current uncertainties. Over the next 12-24 months, increased transparency and public discussion are expected, potentially influencing policy decisions and funding allocations across Europe. Monitoring how the model’s openness, native data use, and strategic objectives evolve will be critical for assessing the future of European AI sovereignty.
Key Questions
What are the main concerns about AMÁLIA’s openness?
It is unclear how accessible the model will be to external researchers and whether its training data and architecture will be fully transparent, which are key factors for assessing its openness and reproducibility.
How much native Portuguese data was used in training AMÁLIA?
Approximately 5.8 billion tokens from Portugal’s national web archive were used in the extended pre-training, but the sufficiency of this data for strategic purposes remains debated.
What are the strategic goals of the AMÁLIA project?
The official goal is to develop a high-performing Portuguese language model, but specific priorities—such as sovereignty, openness, or commercial deployment—have not been publicly clarified.
Will the final version address current uncertainties?
It is expected that the June 2026 release will clarify some questions, but until then, many aspects remain uncertain and subject to ongoing development.
Why do these questions matter for European AI development?
Addressing these questions is vital for ensuring European models are truly sovereign, transparent, and aligned with national strategic interests, rather than being mere replicas of larger, opaque foreign models.
Source: ThorstenMeyerAI.com