Related Experiment Videos
A Benchmark Corpus of Yemeni Proverbs for Figurative and Cultural Language Modeling
Nasser Thmer1,2, Ali Al-Laith3, Muhammad Shoaib4
1Computer Science Department, University of Engineering and Technology Lahore, Lahore, Pakistan. nasserthmer.net@gmail.com.
None:
We present a structured corpus of 5,252 Yemeni Arabic proverbs paired with explanations in Modern Standard Arabic (MSA). The dataset was compiled from four printed proverb anthologies and three publicly accessible online repositories between January and June 2024. All entries were manually transcribed or programmatically extracted and subsequently verified to ensure accuracy and fidelity to the original sources. Each record includes the proverb text, its explanation, source attribution, and available geographic metadata. The corpus addresses the scarcity of dialect-specific Arabic resources and provides a benchmark for evaluating figurative language understanding and culturally grounded language generation. Technical validation through supervised fine-tuning and unsupervised clustering demonstrates the dataset's internal consistency and thematic diversity. The dataset is publicly available via Zenodo under a Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Related Concept Videos
Modeling and Similitude
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Components of Language
Language and Cognition
Typical Model Studies
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...