Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

A review of optimization strategies for deep and machine learning in diabetic macular edema.

Frontiers in artificial intelligence·2026
Same author

Plant health index as an anomaly detection tool for oil refinery processes.

Scientific reports·2022
Same author

Building a planter system using waste materials using value engineering environmental assessment.

Scientific reports·2022
Same author

On the accuracy of ARIMA based prediction of COVID-19 spread.

Results in physics·2021
Same author

A study on the efficiency of the estimation models of COVID-19.

Results in physics·2021

相关实验视频

Updated: Jul 14, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

阿拉伯语语音识别模型使用百度的深度和集群学习.

Fawaz S Al-Anzi1, Bibin Shalini Sundaram Thankaleela1

  • 1Department of Computer Engineering, College of Engineering and Petroleum, Kuwait University, Kuwait.

Frontiers in artificial intelligence
|September 22, 2025
PubMed
概括

这项研究提高了阿拉伯语语音识别,使用K-means集群在Mel-frequency cepstral系数 (MFCCs) 和百度上.

科学领域:

  • 计算语言学 计算语言学
  • 机器学习 机器学习
  • 语音处理 语音处理

背景情况:

  • 阿拉伯语自动语音识别 (ASR) 由于语言复杂性而存在独特的挑战.
  • 在ASR中,无监督学习方法对于处理未标记的音频数据至关重要.
  • 现有的ASR模型需要重要的标记数据来进行有效的培训.

研究的目的:

  • 通过使用无监督集群技术来提高阿拉伯语ASR的准确性.
  • 评估各种分类算法对集群音频特征的性能.
  • 为了证明百度对阿拉伯语ASR的Deep Speech框架的有效性.

主要方法:

  • 从未标记的阿拉伯语音频中提取Mel频率塞普斯特拉系数 (MFCC).
  • 应用K-means集群用于MFCC特征的无监督分组.
  • 使用决策树,XGBoost,KNN和随机森林进行分类;使用集群数据培训和测试百度的深度语言.

主要成果:

  • K-means集群成功分类了声学上相似的阿拉伯语音频段.
  • 百度的Deep Speech模型实现了0.3720的低文字错误率 (WER) 和0.0568.8的字符错误率 (CER).
  • 提出的方法表明,阿拉伯语ASR的准确性显著增加.
关键词:
拜杜斯的演讲是深层次的.一个RNN RNN一个声学模型.聚类集群是指聚类的聚类.深度学习是一种深度学习.语言模型语言模型

更多相关视频

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)
10:55

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)

Published on: April 11, 2026

相关实验视频

Last Updated: Jul 14, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)
10:55

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)

Published on: April 11, 2026

结论:

  • 无监督的MFCC集群是阿拉伯ASR的有效预处理步骤.
  • 百度的Deep Speech框架在集群数据上训练时,可以提供高性能阿拉伯语语音识别.
  • 综合方法提高了阿拉伯语言的ASR准确性,精度和效率.