📊 Full opportunity report: AMÁLIA · The Three Hard Questions. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Portugal launched AMÁLIA, a state-funded European Portuguese LLM, which outperforms many models but raises critical questions about transparency, data sufficiency, and objectives. These issues have broad implications for Europe’s sovereign AI efforts.
Portugal’s €5.5 million investment in the AMÁLIA large language model has resulted in a functioning European Portuguese AI system, but critical questions about its openness, data sufficiency, and strategic focus remain unanswered, raising concerns about the broader European sovereign-LLM movement.
AMÁLIA, developed by a consortium of approximately 60 researchers across Portugal’s leading institutions, was officially launched in October 2025. It is based on a continuation of the EuroLLM multilingual foundation, with the base version completed in September 2025 and currently available to 450,000 academic users through the FCT’s IAedu platform. The model handles text only, with multimodal capabilities planned for future updates. It has been shown to outperform previous open models on Portuguese benchmarks and surpasses Qwen 3-8B on most tests, though it still trails on certain tasks like ALBA.
However, critiques by Duarte O.Carmo and others highlight unresolved issues regarding transparency, native-language data volume, and strategic objectives. The project’s technical approach relies heavily on extending existing multilingual models rather than training from scratch, which raises questions about the model’s native-language proficiency and the adequacy of the data used.
AMÁLIA
The three hard
questions.
Portugal spent €5.5M to build a European Portuguese LLM. The base version is operational, the benchmarks beat Qwen 3-8B on most pt-PT tasks. So why are the most important questions still unanswered?
Last month, Duarte O.Carmo published the sharpest public analysis of AMÁLIA — Portugal’s state-funded European Portuguese large language model. He prefaces his critique with the necessary diplomatic apparatus before doing what almost nobody else in the European-sovereign-LLM discourse has been willing to do publicly: asking hard questions about whether the work, as released, actually does what it set out to do. This piece is a structural extension of his analysis. The AMÁLIA case study exposes three hard questions every national LLM effort needs to answer publicly — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
Three questions every national LLM effort needs to answer publicly.
Duarte O.Carmo’s framing maps cleanly onto the structural argument. Each question lands specifically in AMÁLIA — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
The three questions form a structural feedback loop. Q3 (optimization target) determines Q2 (data volume needed) which conditions Q1 (openness sufficient for community contribution). The European sovereign-LLM movement collectively benefits from these questions becoming standard methodology disclosure, not exceptional critique.
European Portuguese language large language model
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
107 billion tokens. 5.8 billion clearly pt-PT.
The structurally tractable question with a structurally surprising answer. For a model whose entire stated purpose is European Portuguese prioritization, the native-language share of extended pre-training is 5.5%. The implications cascade into every other question.
AI transparency tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Olmo standard. AMÁLIA’s current state.
Allen Institute for AI’s Olmo project defines what “fully open” operationally requires. Olmo doesn’t lead frontier benchmarks. That’s not the point. The point is to be the structural reference for openness. AMÁLIA’s “fully open source” claim should track to the operational standard.
multilingual AI model
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Four strategic positions. AMÁLIA between two and three.
Approximately €100M+ in publicly disclosed European sovereign-LLM funding across the major initiatives. The structural question every project faces: what is the actual competitive position you’re staking? Four options — none mutually exclusive — but each requiring different commitments.
AI data annotation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three standards. For AMÁLIA and the movement.
The structural critique generalizes beyond AMÁLIA. Italy, France, Germany, Switzerland, the OpenEuroLLM consortium, and every subsequent national project benefit from public discourse holding national LLM efforts to operational standards on openness, data accounting, and strategic positioning.
The European sovereign-AI agenda is a serious strategic project that deserves serious public discourse. O.Carmo’s analysis is what serious public discourse looks like. Appropriately diplomatic. Structurally rigorous. Willing to ask the hard questions in public when the public investment justifies it. More of this is needed — across every European sovereign-LLM project, not just AMÁLIA.
Implications for European Sovereign AI Strategies
The development of AMÁLIA exemplifies the broader challenge faced by European countries: balancing transparency, data sufficiency, and strategic goals in national AI projects. The questions raised about openness and native data are not unique to Portugal but are central to the continent’s efforts to develop independent, trustworthy AI systems. How these issues are addressed will influence Europe’s ability to create competitive, sovereign models that meet local language and cultural needs while maintaining transparency and accountability.
European Sovereign-LLM Initiatives Face Common Challenges
Across Europe, multiple nations are pursuing large language models with public funding, including Italy’s Minerva, Germany’s Aleph Alpha, France’s Mistral, and others. These projects often share common technical approaches—such as building on multilingual foundations—and face similar questions about openness, native-language data, and strategic priorities. The European Union has also shown interest in fostering a cohesive framework for sovereign AI development, emphasizing transparency and control.
Portugal’s AMÁLIA is the most publicly significant example due to its substantial investment and national scope. Its progress and the questions it raises are indicative of the broader structural issues facing European sovereign-LLM efforts, which are still in early stages and often lack clear answers to critical questions about data, openness, and objectives.
“AMÁLIA is an impressive piece of work, but its transparency and native-language data sufficiency require serious scrutiny.”
— Duarte O.Carmo
Unanswered Questions About AMÁLIA’s Openness and Data
It remains unclear how open AMÁLIA truly is, especially regarding access to training data and model weights. The extent of native-language data used, beyond the 5.8 billion tokens from Portuguese web archives, is not fully disclosed. Additionally, the strategic priorities guiding the project—whether it aims for transparency, commercial deployment, or academic research—are still under debate. The final version, expected in June 2026, may address some of these gaps, but current information is limited.
Next Steps for AMÁLIA and European Sovereign Models
In the coming months, the AMÁLIA team will release the final version, which may clarify some of the current uncertainties around data and openness. Simultaneously, broader European initiatives are likely to scrutinize and compare approaches, emphasizing transparency and native-language capabilities. Policymakers and researchers will monitor these developments to assess whether these models meet the goals of sovereignty, trustworthiness, and linguistic relevance.
Further technical evaluations, transparency disclosures, and strategic clarifications are expected as the project matures, shaping the future landscape of European AI independence.
Key Questions
What makes AMÁLIA different from other European language models?
AMÁLIA is based on extending a multilingual foundation rather than training from scratch, and it is publicly funded by Portugal, making it a key case in the European effort for sovereign AI. Its performance and development approach are representative of broader European strategies.
Why are questions about openness and native data important?
Transparency about data sources and model access is crucial for trust, accountability, and strategic independence. Without clarity, models risk being perceived as opaque or unreliable, which could undermine their adoption and the continent’s AI sovereignty.
When will the final version of AMÁLIA be available?
The final version is expected in June 2026, which may include further disclosures and improvements based on ongoing evaluations and stakeholder feedback.
How does AMÁLIA compare to models like Qwen 3-8B?
AMÁLIA outperforms Qwen 3-8B on most Portuguese benchmarks but still trails on some tasks like ALBA. Its technical approach and data strategy differ, emphasizing continuation of existing multilingual models rather than training from scratch.
Source: ThorstenMeyerAI.com