Abu Dhabi-based Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) has developed ArabCulture-Dialogue, a benchmark designed to test whether AI models understand cultural context across Modern Standard Arabic and 13 national dialects.
The benchmark was built with 26 native Arabic speakers from 13 countries. Its dataset covers 12 everyday topics, including weddings, food, parenting, agriculture, arts and games, with more than 340,000 words across 54 subtopics.
Researchers tested models on three tasks: selecting culturally appropriate replies, translating between Modern Standard Arabic and a specific dialect, and continuing conversations in a named dialect.
According to the study, leading models identified culturally appropriate responses with accuracy in the mid-90% range. However, they produced the correct target-country dialect in only about half of cases. North African and Emirati dialogues were among the most difficult.
The research evaluated Arabic-focused models including Jais, ALLaM and SILMA, alongside multilingual and proprietary systems. Smaller open-weight Arabic models performed the weakest on dialect-generation tasks, in some cases approaching random guessing.
The findings suggest that models may hold relevant cultural knowledge but struggle to express it naturally in local dialects. Specifying the country and region associated with a conversation improved accuracy, the researchers said. The work was presented at the 63rd Annual Meeting of the Association for Computational Linguistics.
Source: Middle East AI News


