ovr.news

Archaeology, rediscovered knowledge, the past opening up

Arabic LLMs struggle with cultural norms, not facts

arxiv.org · 16 September 2026

Summary and headline written by AI from the source article. How we work

Researchers evaluated how well large language models understand and respond to culturally appropriate prompts in Arabic.

They created AraBehave, a new benchmark of 1,623 open-ended prompts and 29,214 human judgments from speakers across the Arab world. The team found cultural appropriateness splits into two parts: taking the right position on an issue and having accurate cultural knowledge.

While both leading general-purpose models and the best Arabic-focused model achieved similar overall scores (3.84 and 3.83 out of 5), they failed for different reasons. General models often had correct facts but expressed opinions that didn't align with cultural norms, particularly around secular viewpoints and presenting multiple sides of settled issues. The top Arabic model, however, sometimes invented religious traditions or misquoted sacred texts. Importantly, the researchers discovered that improving a model’s cultural stance is relatively easy.

A single sentence of cultural instruction boosted the score of one model to 4.57, exceeding all Arabic-specialized models. Conversely, asking a model to be “clear and objective” lowered the score of one model by 0.68 points. Cultural understanding appears to rely on both model size and training on Arabic data, unlike standard safety benchmarks which show little variation. The AraBehave benchmark, annotations, and scoring model will be publicly released.

Was this worth your time?
Read on arxiv.org
Surfaced by the Discovery lens — one of the vital signs ovr.news reads.
How we evaluated this

More in Discovery

Browse all Discovery articles

What made it worth it?

What's wrong with this article?