Related Experiment Video
Updated: Aug 19, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Augmented-syllabification of n-gram tagger for Indonesian words and named-entities
Suyanto Suyanto1, Andi Sunyoto2, Rezza Nafi Ismail1
1School of Computing, Telkom University, Bandung, Indonesia.
Abstract:
As one of the statistical-based models, an n-gram syllabification commonly gives a high syllable error rate (SER) for Bahasa Indonesia, one of the low-resource languages, since it fails for a high out-of-vocabulary (OOV) rate. Two previous models: bigram-syllabification with flipping onsets (BFO) and a combination of bigram with backoff smoothing based on phonological similarity (CBSPS), which use augmentation methods, can reduce the OOV rate. However, there are two problems in both BFO and CBSPS. First, they use an n-gram that is applied syllable-level, instead of grapheme-level, so that they suffer on the sparsity of n-grams. Second, they rely on a procedure to detect the positions of both vowels and diphthongs. Both problems make them not capable of distinguishing diphthongs from derivative words as well as syllabifying named-entities, which have many ambiguities related to vowels and semi-vowels. In this paper, a syllabification based on an n-gram tagger, which is applied on grapheme-level and does not rely on both vowel and diphthong detections, is developed to solve both problems. Besides, three data augmentation methods are exploited to enrich the dataset. The 5-fold cross-validations (5-FCV) using both datasets of 50 k words and 15 k named-entities show that the proposed augmented-syllabification of n-gram tagger (ASnGT) model is significantly better than both BFO and CBSPS. It is also significantly better than the fuzzy k-nearest neighbor in every class (FkNNC)-based model for formal words and named-entities. However, it suffers from derivative words, where it cannot easily distinguish them from both absorption words and terms of foreign languages. Besides, it also undergoes some foreign named-entities.
More Related Videos
Related Concept Videos
Tagging and Fusion Proteins
Air-entraining Agents
Nomenclature of Alkanes
The alkane nomenclature considers the length of the carbon chain, the number, and the location of the substituent to arrive at its systematic name. The IUPAC...
Mnemonic Devices
Acronyms
Acronyms are created by using the initial letters of a series of words to form a new word or phrase. This approach condenses complex information into a single, memorable entity. For example,...
Tip-of-the-Tongue Phenomenon
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...

