Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

1.3K
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
1.3K
Auditory Perception01:17

Auditory Perception

1.4K
The auditory system is essential for sound perception, utilizing various critical structures. When sound waves enter the outer ear, they travel through the ear canal and cause the eardrum to vibrate. These vibrations are then transmitted to the middle ear, where three tiny bones – the malleus, incus, and stapes – amplify the sound. This amplification is crucial, as it ensures that the sound vibrations are strong enough to be conveyed to the inner ear. These vibrations then reach the...
1.4K
Auditory Pathway01:15

Auditory Pathway

8.8K
Auditory pathways constitute the complex neural circuits responsible for transmitting and interpreting auditory information from the peripheral auditory system to the brain. Sound waves are initially captured by the outer ear, funneled through the ear canal, and reach the tympanic membrane (eardrum). These vibrations are transmitted via the middle ear's ossicles to the inner ear's cochlea.
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
8.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Preoperative prediction of occult high-volume central lymph node metastasis in clinically node-negative papillary thyroid carcinoma using a multimodal practical model: development, temporal validation, and external validation.

Diagnostic and interventional radiology (Ankara, Turkey)·2026
Same author

Noninvasive detection and prediction method for temperature in <i>ex vivo</i> biological tissue based on opto-thermal-acoustic co-coupling.

Journal of biomedical optics·2026
Same author

FDPMambaFuse: A frequency-domain and parallel Mamba-based model for multimodal medical image fusion.

Neural networks : the official journal of the International Neural Network Society·2026
Same author

A self-assembled HSA-ICG nanoprobe enables accurate photoacoustic thermometry and feedback-controlled tumor photothermal therapy.

Biomaterials science·2026
Same author

Quantitative photoacoustic imaging algorithm using sparse decomposition for photoacoustic and ultrasound dual-mode imaging.

Biomedical optics express·2026
Same author

Photodynamic Therapeutic Monitoring of Glioblastoma Using High-Resolution Photoacoustic Vascular Imaging.

Molecular pharmaceutics·2026

Related Experiment Video

Updated: Mar 27, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.1K

A multimodal attention fusion-based model for pathological voice detection.

Baiya Li1, Boheng Zhang2,3,4, Haorui Huang2,3

  • 1Department of Otolaryngology Head and Neck Surgery, The First Affiliated Hospital of Xi'an Jiaotong University, Xi'an, People's Republic of China.

Biomedical Physics & Engineering Express
|March 25, 2026
PubMed
Summary

This study introduces a Multimodal Fusion Network (MFNet) for improved pathological voice classification. MFNet effectively distinguishes multiple laryngeal disorders using combined voice signal features, aiding early screening.

Keywords:
deep learningmultimodal fusionpathological voice detectiontunable wavelet transform

Related Experiment Videos

Last Updated: Mar 27, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.1K

Area of Science:

  • Speech processing
  • Biomedical engineering
  • Machine learning for healthcare

Background:

  • Pathological voice detection offers non-invasive screening for laryngeal disorders.
  • Current methods often use single acoustic features, limiting fine-grained classification of multiple voice pathologies.

Purpose of the Study:

  • To propose an advanced Multimodal Fusion Network (MFNet) for accurate pathological voice classification.
  • To enhance the discrimination of multiple pathological voice categories beyond binary classification.

Main Methods:

  • Developed MFNet integrating time-domain (SincNet) and frequency-domain (TQWT, SKNet) features.
  • Employed an Attentional Feature Fusion (AFF) module for adaptive feature integration.
  • Utilized raw speech waveforms and Mel-spectrograms for comprehensive analysis.

Main Results:

  • MFNet outperformed existing baseline models on three public pathological voice datasets.
  • Demonstrated strong robustness and generalization capabilities across diverse datasets.
  • Ablation studies confirmed the effectiveness of the proposed architectural components.

Conclusions:

  • MFNet offers an effective solution for multi-category pathological voice classification.
  • The proposed method shows significant potential for computer-aided screening of voice disorders.
  • Multimodal feature fusion enhances the discriminative power for laryngeal disorder detection.