Related Experiment Videos
A large-scale annotated Urdu corpus and deep learning benchmark for mental health classification.
Maleeka Fatima1, Muhammad Saleem Khan1, Muhammad Shahzad Faisal1
1Department of Computer Science, COMSATS University Islamabad, Attock, 43600, Pakistan.
Scientific Reports
|June 8, 2026
Summary
This study introduces a deep learning framework for detecting anxiety and depression in Urdu social media text. UrduBERT achieved the highest accuracy, enabling mental health monitoring in underrepresented linguistic communities.
Area of Science:
- Computational linguistics
- Mental health informatics
- Artificial intelligence in healthcare
Background:
- Mental health disorders like anxiety and depression are a global concern, necessitating early detection.
- Traditional diagnostic methods face accessibility and stigma challenges.
- AI and NLP show promise for automated mental health detection via social media text, but research is limited in low-resource languages.
Purpose of the Study:
- To develop and evaluate a deep learning framework for mental health classification (anxiety, depression, neutral) in Urdu.
- To create the first large-scale, publicly available Urdu dataset for mental health classification.
- To establish performance benchmarks for Urdu mental health detection using AI.
Main Methods:
- Creation of a novel annotated dataset of 36,000 Urdu tweets, classified into anxiety, depression, and neutral categories.
- Systematic translation, manual annotation, and automated labeling with rigorous quality validation.
- Adaptation and evaluation of three deep learning architectures: CNN+BiLSTM, CNN+BiGRU, and UrduBERT.
Main Results:
- UrduBERT achieved the highest accuracy (81.71%), outperforming CNN+BiLSTM (79.08%) and CNN+BiGRU (78.25%).
- UrduBERT demonstrated superior precision, recall, and F1-scores across all classification categories.
- Transformer-based models like UrduBERT excel in capturing linguistic nuances in morphologically rich languages like Urdu.
Conclusions:
- The proposed deep learning framework, particularly UrduBERT, offers a robust solution for mental health classification in Urdu.
- This work bridges a critical gap in computational mental health research for low-resource languages.
- The framework supports scalable, culturally sensitive mental health monitoring and early detection in underrepresented communities.