Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Long-term Depression01:03

Long-term Depression

3.1K
Long-term depression, or LTD, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTD is the process of synaptic weakening that occurs over time between pre and postsynaptic neuronal connections. The synaptic weakening of LTD works in opposition to synaptic strengthening by long-term potentiation (LTP) and together are the main mechanisms that underlie learning and memory.
Calcium Ion Concentration Mechanism
If over...
3.1K
Long-term Depression01:05

Long-term Depression

33.1K
Long-term depression, or LTD, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTD is the process of synaptic weakening that occurs over time between pre and postsynaptic neuronal connections. The synaptic weakening of LTD works in opposition to synaptic strengthening by long-term potentiation (LTP) and together are the main mechanisms that underlie learning and memory.
33.1K
Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

925
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
925
Depressive Disorders: Etiology01:27

Depressive Disorders: Etiology

433
Depressive disorders result from a complex interplay of biological, psychological, and sociocultural factors, each contributing uniquely to the development and persistence of the condition. Understanding these factors provides critical insight into the multifaceted nature of depression.
Biological Factors in Depression
Biological predispositions significantly influence the risk of developing depressive disorders. Genetic studies highlight the role of variations in the serotonin transporter...
433
Frequency-Domain Interpretation of PD Control01:24

Frequency-Domain Interpretation of PD Control

346
Proportional-Derivative (PD) controllers are widely used in fan control systems to improve stability and performance. A fan control system can be effectively represented using a Bode plot to illustrate the impact of a PD controller through its transfer function. The Bode plot visually conveys how PD control modifies the fan's response across various frequencies, providing a frequency domain interpretation of the controller's behavior.
The proportional control gain, combined with the...
346
Depressive Disorders: MDD and Dysthymia01:27

Depressive Disorders: MDD and Dysthymia

618
Depressive disorders are a group of mental health conditions characterized by pervasive feelings of sadness, diminished pleasure in life, and a significant impact on daily functioning. These conditions are most prevalent in individuals during their 30s and affect women at twice the rate of men. Contrary to popular belief, younger individuals are generally more susceptible to these disorders than older adults. Two key types of depressive disorders include Major Depressive Disorder (MDD) and...
618

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Parkinson's disease classification using optimized attention-based deep learning from EEG signals with interpretable sub-band topography.

Brain informatics·2026
Same author

Dual-stream token fusion with Swin Transformer and lesion-aware tokens for gastric metaplasia classification in IoMT-assisted deployment.

BMC medical imaging·2026
Same author

Clinically Deployable Handwriting Biomarkers of Parkinson's Disease via Multiscale Attention and Bayesian-Genetic Optimization.

Brain and behavior·2026
Same author

EEG-based schizophrenia detection using handcrafted biomarkers and a TOA-optimized hybrid multi-branch CNN-Transformer framework.

Brain research bulletin·2026
Same author

Intelligent <i>in-silico</i> prioritization of antimalarial peptide candidates under explicit physicochemical windows via <i>de novo</i> CTCM-Neo generation and conformal-gated calibrated classification.

Frontiers in cellular and infection microbiology·2026
Same author

Optimized CNN-based facial analysis for depression detection: Managing mental disorder in education.

Brain research bulletin·2026

Related Experiment Video

Updated: Jan 13, 2026

Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
08:45

Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example

Published on: October 24, 2012

15.2K

Depression detection from speech data using deep learning-based optimized temporal-frequency-channel attention with

Khosro Rezaee1

  • 1Department of Biomedical Engineering, Meybod University, Meybod, Iran.

Journal of Affective Disorders
|January 8, 2026
PubMed
Summary

This study introduces a novel deep learning framework for detecting depression from voice recordings. The interpretable model achieves high accuracy across datasets, offering a promising tool for mental health screening.

Keywords:
Attention mechanismsComputer-assisted diagnosisDeep learningDepression detectionOptimization algorithmSpeech acoustics

More Related Videos

Author Spotlight: Therapeutic Benefit of Closed-Loop Deep Brain Stimulation in Depression Treatment
05:19

Author Spotlight: Therapeutic Benefit of Closed-Loop Deep Brain Stimulation in Depression Treatment

Published on: July 7, 2023

3.2K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

807

Related Experiment Videos

Last Updated: Jan 13, 2026

Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
08:45

Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example

Published on: October 24, 2012

15.2K
Author Spotlight: Therapeutic Benefit of Closed-Loop Deep Brain Stimulation in Depression Treatment
05:19

Author Spotlight: Therapeutic Benefit of Closed-Loop Deep Brain Stimulation in Depression Treatment

Published on: July 7, 2023

3.2K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

807

Area of Science:

  • Artificial Intelligence
  • Speech Signal Processing
  • Clinical Psychology

Background:

  • Detecting depression from voice is challenging due to subtle acoustic cues and poor cross-lingual generalization of current models.
  • Existing methods often require transcriptions or visual data, limiting their applicability.
  • Individual variations in speech patterns complicate accurate depression detection.

Purpose of the Study:

  • To develop a lightweight, interpretable deep learning framework for depression detection directly from raw speech audio.
  • To overcome limitations of cross-lingual generalization and reliance on transcriptions.
  • To enhance model robustness and adaptability to diverse acoustic conditions.

Main Methods:

  • A streamlined ResNet-18 model enhanced with a Temporal-Frequency-Channel Attention (TFCA) unit processes speech spectrograms.
  • Raw audio is segmented into clips and converted to time-frequency representations.
  • A novel Parameter Optimization with Conscious Allocation using Iterative Intelligence (POCAII) strategy optimizes hyperparameters for faster convergence and robustness.

Main Results:

  • Achieved 89.38% segment-level accuracy and 93.94% subject-level accuracy on the DAIC-WOZ dataset.
  • Reached 89.96% segment-level accuracy and 93.23% subject-level accuracy on the Androids Corpus.
  • Demonstrated high segment-level Area Under the Receiver Operating Characteristic Curve (AUC) of 95.3% and 95.7% respectively, with interpretable attention visualizations.

Conclusions:

  • The proposed deep learning framework effectively detects depression from speech audio with high accuracy and interpretability.
  • The model shows strong cross-lingual and cross-dataset generalization capabilities.
  • This approach offers a promising, transcription-free solution for scalable depression screening.