Vietnamese Minority Languages Get First NLP Dataset

Researchers at the VinUniversity in Hanoi, Vietnam, created a new dataset to help computers better understand Cham, Khmer, and Tay-Nung, three distinct languages spoken by ethnic minorities within the country. These languages have been largely excluded from Natural Language Processing research due to limited data and unique linguistic challenges.
The corpus, called CKTN, contains 44,367 documents and 24 million pieces of text. The team found that existing multilingual language models struggle to accurately process these languages, often failing to grasp semantic meaning despite appearing proficient in basic tasks.
Standard methods for adapting these models can actually reinforce errors due to differences in writing systems and varying degrees of influence from the Vietnamese language. To address this, the researchers developed a new adaptation technique focusing on script awareness.
This improves performance on tasks like text classification and highlights the inadequacy of simple retrieval methods for evaluating true language understanding. Further work will focus on refining these models and expanding the dataset.
Surfaced by the Discovery lens — one of the vital signs ovr.news reads.
How we evaluated this
AI summary
read the original for the full story — Read on arxiv.org . How we work →