Related Experiment Videos
Morpheme matching based text tokenization for a scarce resourced language
Zobia Rehman1, Waqas Anwar, Usama Ijaz Bajwa
1Department of Computer Science, COMSATS Institute of Information Technology, Abbottabad, Pakistan.
Abstract:
Text tokenization is a fundamental pre-processing step for almost all the information processing applications. This task is nontrivial for the scarce resourced languages such as Urdu, as there is inconsistent use of space between words. In this paper a morpheme matching based approach has been proposed for Urdu text tokenization, along with some other algorithms to solve the additional issues of boundary detection of compound words, affixation, reduplication, names and abbreviations. This study resulted into 97.28% precision, 93.71% recall, and 95.46% F1-measure; while tokenizing a corpus of 57000 words by using a morpheme list with 6400 entries.
Related Concept Videos
Components of Language
Sign Test for Matched Pairs
To conduct the sign test, we first calculate the differences in value between...
Mismatch Repair
Mismatch Repair
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
Pre-mRNA Processing: RNA Splicing
Alternative RNA Splicing
There are five types of alternative RNA splicing that vary in the ways the pre-mRNA segments are removed or retained in the mature mRNA. The first...