AI benchmark bridges Arabic-Russian science knowledge
Summary and headline written by AI from the source article. How we work
Researchers developed a new benchmark and parallel corpus to improve scientific translation between Arabic and Russian, aiming to enhance global collaboration.
The work addresses language barriers that hinder research exchange between Arabic-speaking and Russian-speaking scientific communities. They compiled a hybrid parallel corpus containing approximately 27,000 sentence pairs from scientific abstracts and general texts. This resource was used to fine-tune three multilingual language models: mT5-base, NLLB-200-distilled-1.3B, and Qwen2.5-7B-Instruct.
The Qwen2.5-7B model, fine-tuned with a specific method called QLoRA, achieved the best performance metrics, showing substantial improvements over a basic zero-shot approach. The research found that few-shot prompting did not boost performance, indicating the necessity of domain-specific fine-tuning for effective translation. The team released the corpus, models, and evaluation code to facilitate knowledge transfer and support sustainable scientific partnerships.


