Better Language Docs: POS Tagging Boosts Automated Translation

Researchers at the University of Hawaiʻi at Mānoa developed a machine-learning pipeline to speed up the creation of interlinear glosses, detailed translations that aid in documenting endangered languages, specifically for Irabu, a Southern Ryukyuan language spoken in Japan. Creating these glosses is time-consuming, often taking an hour of work for each minute of recorded speech.
The team found that including part-of-speech (POS) tagging as part of the annotation process significantly improves the accuracy of automated glossing. Adding POS information improved glossing accuracy by 4.4 points, and the benefit increased when less training data was available.
Currently, the automated POS tagger makes errors on about 12% of morphemes, but improving its accuracy could unlock even greater gains in efficiency. The researchers recommend annotating texts with four layers, text, POS tags, glosses, and translations, to maximize the value of automated tools.
Surfaced by the Discovery lens — one of the vital signs ovr.news reads.
How we evaluated this
AI summary
read the original for the full story — Read on arxiv.org . How we work →