Related Experiment Video
Updated: Jun 22, 2025

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
A comprehensive dataset for Arabic word sense disambiguation
1Computing and Applied Technology, College of Technological Innovation, Zayed University, UAE.
This study introduces a new Arabic dataset for word sense disambiguation, covering 100 polysemous words and 367 senses. It aids Arabic Natural Language Processing (NLP) by including rare senses and diverse contexts.
Area of Science:
- Computational Linguistics
- Natural Language Processing
- Artificial Intelligence
Background:
- Arabic language exhibits significant polysemy, posing challenges for NLP tasks.
- Existing datasets may lack comprehensive coverage of diverse word senses, especially rare ones.
- Accurate word sense disambiguation is crucial for advancing Arabic NLP applications.
Purpose of the Study:
- To introduce a novel, comprehensive dataset for Arabic word sense disambiguation.
- To provide a resource that mitigates bias by including all senses, including rare ones.
- To support the development of more robust Arabic NLP models.
Main Methods:
- Collected data from diverse web sources (news, medicine, finance).
- Included 10 contextual sentences per word sense, totaling 3670 samples for 100 polysemous words.
- Utilized GPT3.5-turbo for generating synthetic data for underrepresented rare senses.
Main Results:
- Developed a dataset with 367 unique senses for 100 Arabic polysemous words.
- Dataset spans multiple domains and includes both real-world and synthetically generated data.
- Ensures coverage of all word senses, including infrequently used ones, to reduce model bias.
Conclusions:
- The dataset is a valuable resource for Arabic word sense disambiguation and NLP research.
- Its comprehensive nature and inclusion of rare senses enhance model training and reduce bias.
- The dataset has potential applications beyond disambiguation, including sentiment analysis.
More Related Videos
12:49Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
Nomenclature of Aromatic Compounds with Multiple Substituents
For disubstituted benzene derivatives, with two groups attached to the benzene ring, three constitutional isomers are possible. For example, consider dimethyl benzene, often called xylene, where the second methyl group can be substituted at the second, third, or fourth carbon. The relative position of the substituents is represented by prefixes ortho, meta, or...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Regional Terms
Primarily, the human body has two major regions, the axial and appendicular regions. The axial region comprises regions from the head to the abdomen and makes up the central body axis. In contrast,...
Genome Annotation and Assembly
Nomenclature of Aromatic Compounds with a Single Substituent
Aromatic Compounds: Overview
In 1825, Faraday...