Portugal’s AMALIA Model Shows Promise, and Limits, as a Research Tool

Researchers evaluated AMALIA, Portugal’s publicly funded 9 billion parameter language model, to determine if it accurately measures societal values. The team assessed AMALIA’s ability to code the moral foundation of authority in European Portuguese, comparing its performance to that of larger, open-source models and human coders. AMALIA achieved agreement with human coders within six points of models eight to thirteen times its size.
However, the study revealed that only about half of AMALIA’s coding performance on authority could be directly attributed to the underlying theory being tested. A larger multilingual language model performed better in closing this “recovery gap,” indicating the limitation may reside in the annotator model rather than the translated text itself.
The researchers suggest that while public ownership and linguistic specialization build initial trust in national language models, rigorous calibration is necessary to establish their true value as scientific instruments. They propose a portable audit method to assess epistemic trustworthiness across different models, languages, and tasks.
Surfaced by the Discovery lens — one of the vital signs ovr.news reads.
How we evaluated this
AI summary
read the original for the full story — Read on arxiv.org . How we work →