Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

208
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
208
Chunking and Rehearsal in Sensory Memory01:22

Chunking and Rehearsal in Sensory Memory

202
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of...
202

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Adaptive regularized spectral reduction for stabilizing ill-conditioned bone-conducted speech signals.

PeerJ. Computer science·2025
Same author

TranStutter: A Convolution-Free Transformer-Based Deep Learning Method to Classify Stuttered Speech Using 2D Mel-Spectrogram Visualization and Attention-Based Feature Representation.

Sensors (Basel, Switzerland)·2023
Same author

Data augmentation and deep neural networks for the classification of Pakistani racial speakers recognition.

PeerJ. Computer science·2022
Same author

Multi-label emotion classification of Urdu tweets.

PeerJ. Computer science·2022
Same author

Comparative analysis on Facebook post interaction using DNN, ELM and LSTM.

PloS one·2019
Same author

A Database as a Service for the Healthcare System to Store Physiological Signal Data.

PloS one·2016

相关实验视频

Updated: Jun 27, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

443

基于Whisper细分的实时多语言语音识别和扬声器日记系统.

Ke-Ming Lyu1, Ren-Yuan Lyu1, Hsien-Tsung Chang1,2,3

  • 1Computer Science and Information Engineering, Chang Gung University, Taoyuan, Taiwan.

PeerJ. Computer science
|April 25, 2024
PubMed
概括

本研究介绍了一个实时的多语言语音系统,使用OpenAI的Whisper模型进行准确的语音日记化 (SD) 和自动语音识别 (ASR),即使有口音和语音变化.

关键词:
自动语音识别自动语音识别增量聚类是指增量聚类.实时系统实时系统.演讲者日记化 演讲者日记化

更多相关视频

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K
Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
08:32

Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks

Published on: September 5, 2019

5.6K

相关实验视频

Last Updated: Jun 27, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

443
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K
Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
08:32

Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks

Published on: September 5, 2019

5.6K

科学领域:

  • 语音处理 语音处理
  • 人工智能的人工智能
  • 自然语言处理自然语言处理.

背景情况:

  • 传统的语音识别和说话者日记化系统在动态的多语言环境中扎,尤其是口音和频繁的说话者变化.
  • 在当前的系统中,准确处理普通话语的语音与台湾口音以及管理发言者的转换是重大挑战.

研究的目的:

  • 开发一套尖端的,实时的多语言语音识别和演讲者日记系统.
  • 在复杂的多音箱场景中提高性能,重点是用台湾口音和快速音箱转换的普通话.

主要方法:

  • 利用 OpenAI 的 Whisper 模型来实现核心语音识别功能.
  • 集成先进的扬声器日记化技术,包括高效的输出处理和扬声器嵌入技术,用于实时应用.
  • 优化了系统的动态,多扬声器环境与频繁的扬声器开关.

主要成果:

  • 实现了一个有前途的单词日记化错误率 (WDER) 总体为6.96%.
  • 在双扬声器 (2.68% WDER) 和三扬声器 (11.65% WDER) 场景中表现出强的性能.
  • 实时性能与非实时基线模型可比.

结论:

  • 开发的系统代表了实时多语言语音处理的重大进步.
  • 该系统有效地处理复杂的对话动态,包括不同的口音和多个扬声器.
  • 这项研究验证了将先进的ASR和SD技术集成到现实应用中的有效性.