ovr.news

Archaeology, rediscovered knowledge, the past opening up

Better AI for African Languages Boosts Accuracy

arxiv.org · 17 July 2026
Better AI for African Languages Boosts Accuracy
Photo: arxiv.org
Read on arxiv.org

Researchers at an unspecified institution developed VEXMLM, a language model that improves performance for Amharic and Tigrinya, two Ge'ez-script languages, and 17 other low-resource African languages. Existing multilingual AI models struggle with these languages due to limited vocabulary and excessive fragmentation of words into smaller units.

The team trained a new tokenizer specifically for Amharic and Tigrinya, adding 30,000 Ge'ez-script subwords to the XLM-R model. They then trained VEXMLM in two stages, first with masked language modeling and then with supervised fine-tuning on tasks like question answering and sentiment analysis.

The improved model outperformed XLM-R and Glot500 on these tasks. Notably, VEXMLM increased accuracy in identifying named entities, even in languages not directly used in its training, by reducing the rate of out-of-vocabulary tokens. The researchers suggest this approach could help bridge the gap in AI capabilities for underrepresented languages.

Surfaced by the Discovery lens — one of the vital signs ovr.news reads.

How we evaluated this
AI summary

read the original for the full story — Read on arxiv.org . How we work →

Why are you reporting this article?

Why are you reporting this article?