Better AI for African Languages Boosts Accuracy

Researchers at an unspecified institution developed VEXMLM, a language model that improves performance for Amharic and Tigrinya, two Ge'ez-script languages, and 17 other low-resource African languages. Existing multilingual AI models struggle with these languages due to limited vocabulary and excessive fragmentation of words into smaller units.
The team trained a new tokenizer specifically for Amharic and Tigrinya, adding 30,000 Ge'ez-script subwords to the XLM-R model. They then trained VEXMLM in two stages, first with masked language modeling and then with supervised fine-tuning on tasks like question answering and sentiment analysis.
The improved model outperformed XLM-R and Glot500 on these tasks. Notably, VEXMLM increased accuracy in identifying named entities, even in languages not directly used in its training, by reducing the rate of out-of-vocabulary tokens. The researchers suggest this approach could help bridge the gap in AI capabilities for underrepresented languages.
Surfaced by the Discovery lens — one of the vital signs ovr.news reads.
How we evaluated this
AI summary
read the original for the full story — Read on arxiv.org . How we work →