Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A review of optimization strategies for deep and machine learning in diabetic macular edema.

Frontiers in artificial intelligence·2026
Same author

Plant health index as an anomaly detection tool for oil refinery processes.

Scientific reports·2022
Same author

Building a planter system using waste materials using value engineering environmental assessment.

Scientific reports·2022
Same author

On the accuracy of ARIMA based prediction of COVID-19 spread.

Results in physics·2021
Same author

A study on the efficiency of the estimation models of COVID-19.

Results in physics·2021

Related Experiment Video

Updated: Jul 14, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

Arabic speech recognition model using Baidu's deep and cluster learning.

Fawaz S Al-Anzi1, Bibin Shalini Sundaram Thankaleela1

  • 1Department of Computer Engineering, College of Engineering and Petroleum, Kuwait University, Kuwait.

Frontiers in Artificial Intelligence
|September 22, 2025
PubMed
Summary

This study enhances Arabic speech recognition using K-means clustering on Mel-frequency cepstral coefficients (MFCCs) and Baidu

Keywords:
Baidus deep speechRNNacoustic modelclusteringdeep learninglanguage model

More Related Videos

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)
10:55

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)

Published on: April 11, 2026

Related Experiment Videos

Last Updated: Jul 14, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)
10:55

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)

Published on: April 11, 2026

Area of Science:

  • Computational Linguistics
  • Machine Learning
  • Speech Processing

Background:

  • Arabic Automatic Speech Recognition (ASR) presents unique challenges due to linguistic complexities.
  • Unsupervised learning methods are crucial for handling unlabeled audio data in ASR.
  • Existing ASR models require significant labeled data for effective training.

Purpose of the Study:

  • To improve Arabic ASR accuracy by employing unsupervised clustering techniques.
  • To evaluate the performance of various classification algorithms on clustered audio features.
  • To demonstrate the effectiveness of Baidu's Deep Speech framework for Arabic ASR.

Main Methods:

  • Extraction of Mel-frequency cepstral coefficients (MFCCs) from unlabeled Arabic audio.
  • Application of K-means clustering for unsupervised grouping of MFCC features.
  • Classification using Decision Tree, XGBoost, KNN, and Random Forest; training and testing Baidu's Deep Speech with clustered data.

Main Results:

  • K-means clustering successfully categorized acoustically similar Arabic audio segments.
  • Baidu's Deep Speech model achieved a low word error rate (WER) of 0.3720 and character error rate (CER) of 0.0568.
  • The proposed method demonstrated a significant increase in the accuracy of Arabic ASR.

Conclusions:

  • Unsupervised clustering of MFCCs is an effective preprocessing step for Arabic ASR.
  • Baidu's Deep Speech framework, when trained on clustered data, offers high-performance Arabic speech recognition.
  • The integrated approach enhances ASR accuracy, precision, and efficiency for the Arabic language.