Related Experiment Videos
AryWiki-Instruct: A high-fidelity instruction tuning dataset for Moroccan Arabic (Darija)
Safouane Boudakkou1, Abdelaaziz El Hibaoui1
1Abdelmalek Essaadi University, Faculty of Science, Tetuan, Morocco.
None:
This article presents AryWiki-Instruct, a high-fidelity instruction tuning dataset for Moroccan Arabic (Darija), comprising 46,590 Question and Answer (QA) pairs. The dataset was derived from a snapshot of the Moroccan Arabic Wikipedia (arywiki) and generated using the Gemini-2.5-Flash model via a Context Aware batch processing architecture. The data creation process involved parsing raw XML Wikipedia dumps, filtering for script consistency and token density, and applying a Context Injection generation strategy to prevent coreference ambiguity. To ensure high information density, the raw generated output (82,500 pairs) was subjected to a rigorous automated quality assurance pipeline. This pipeline utilized a hierarchical trigger confirmation algorithm to remove repetitive administrative census noise and employed composite embedding based clustering (multilingual-e5-large) for semantic deduplication. The final dataset is formatted as a JSONL file, providing paired instructions and responses alongside their taxonomic categories. This dataset provides a native first culturally grounded resource designed to facilitate the supervised fine tuning of Large Language Models (LLMs) in Maghrebi dialects, circumventing the syntactic limitations of translated English instruction sets.
Related Concept Videos
Air-entraining Agents
Instrument Calibration
Analytical Balance Calibration
An analytical balance measures mass and requires regular calibration to...
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Reconstruction of Signal using Interpolation
Introduction to MATLAB