Related Experiment Video
Updated: Mar 10, 2026

Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
Published on: June 29, 2021
Heterophonic speech recognition using composite phones.
Ashraf Alkhairy1, Afshan Jafri2
1King Abdul Aziz City for Science and Technology, Riyadh, Saudi Arabia.
This study introduces Composite Phonemes (CP) to improve automatic speech recognition (ASR) for heterophonic languages like Arabic. Composite Phonemes significantly reduce word error rates compared to traditional grapheme-based approaches.
Area of Science:
- Computational Linguistics
- Speech Recognition Technology
- Natural Language Processing
Background:
- Heterophones, words with identical spelling but varied pronunciation, present significant challenges in training Automatic Speech Recognition (ASR) systems.
- Current ASR methods struggle with the inherent ambiguity of heterophonic languages, particularly Arabic, due to orthographic representations lacking phonetic detail.
Purpose of the Study:
- To introduce and evaluate a novel pronunciation unit, the Composite Phoneme (CP), designed to address the challenges posed by heterophonic languages in ASR.
- To develop algorithms for generating CP pronunciations from Modern Orthography (MO) and compare their performance against traditional methods.
Main Methods:
- Developed Composite Phonemes (CP) as a set of alternative phoneme sequences, focusing on consonant-centric units that incorporate short vowels and gemination absent in Modern Orthography (MO).
- Created algorithms to generate CP pronunciations from MO and Simple Phoneme (SP) pronunciations from Classical Orthography (CO).
- Investigated and compared the performance of ASR systems using CP, SP, Undiacritized Grapheme (UG), and Diacritized Grapheme (DG) on the A-SpeechDB corpus.
Main Results:
- ASR systems utilizing CP and SP demonstrated superior performance over UG and DG approaches.
- For an 8,000-word MO vocabulary, CP achieved a Word Error Rate (WER) of 11.78%, outperforming SP_M (12.64%) and SP_A (13.59%).
- For a 24,000-word MO vocabulary, CP's WER was 13.69%, compared to SP_M's 15.08% and SP_A's 16.86%.
Conclusions:
- Composite Phonemes (CP) offer a more effective approach for ASR in heterophonic languages compared to grapheme-based methods.
- While SP shows better performance with uniform statistical models, CP excels in context-independent phone scenarios.
- The CP approach alleviates the necessity for diacritization, simplifying the ASR training process for Arabic.
More Related Videos
Related Concept Videos
Double Resonance Techniques: Overview
Spin decoupling is usually achieved by...
Resonance and Hybrid Structures
Resonance Structures and Resonance Hybrids
The Lewis structure of a nitrite anion (NO2−) may actually be drawn in two different ways, distinguished by the locations of the N–O and N=O bonds.
Components of Language
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
¹H NMR Chemical Shift Equivalence: Homotopic and Heterotopic Protons
Chunking and Rehearsal in Sensory Memory

