Related Experiment Video
Updated: Jun 19, 2025

06:04
Functional Near-Infrared Spectroscopy Hyperscanning Study in Psychological Counseling
Published on: January 17, 2025
477
Using natural language processing to facilitate the harmonisation of mental health questionnaires: a validation study
Eoin McElroy1, Thomas Wood2, Raymond Bond3
1School of Psychology, Ulster University, Coleraine, UK. e.mcelroy@ulster.ac.uk.
BMC Psychiatry
|July 25, 2024
Summary
Natural language processing (NLP) can harmonize mental health questionnaires by matching questions based on meaning. This approach aids cross-study data pooling by identifying semantically similar items, advancing mental health research.
Area of Science:
- Computational linguistics
- Mental health research
- Psychometrics
Background:
- Pooling data from diverse sources enhances mental health research through larger sample sizes and cross-study comparisons.
- Data heterogeneity in variable measurement across studies presents a significant challenge to effective data pooling.
- Natural Language Processing (NLP) offers a potential solution for harmonizing disparate data sources in mental health research.
Purpose of the Study:
- To explore the use of NLP for harmonizing mental health questionnaires by matching questions based on semantic content.
- To assess the accuracy of NLP-derived semantic similarity scores in predicting real-world correlations between questionnaire items.
- To investigate the utility of NLP in identifying semantically similar items for cross-study data pooling.
Main Methods:
- Utilized the Sentence-BERT model to compute semantic similarity (cosine index) for 741 question pairs from five mental health questionnaires.
- Calculated Spearman rank correlations for corresponding item pairs using data from a UK adult sample (N=2,058).
- Estimated the correlation between NLP-derived cosine values and empirical Spearman coefficients, employing network analysis for data structure exploration.
Main Results:
- A moderate correlation (r=0.48, p<0.001) was observed between semantic similarity (cosine index) and empirical correlations (Spearman coefficient).
- NLP-derived cosine scores demonstrated a small error margin in predicting real-world correlations in a holdout sample (MAE=0.05, MedAE=0.04, RMSE=0.064).
- The NLP model identified complex data patterns but required manual rules for network edge inclusion.
Conclusions:
- Quantified semantic similarity between questionnaire items using NLP, showing correlation with empirical responses.
- Demonstrated the potential of NLP to facilitate cross-study data pooling in mental health research by identifying semantically similar items.
- Recommended researchers verify the psychometric equivalence of NLP-matched items before pooling data.
Related Concept Videos
Data Validation
5.0K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
5.0K
Diagnostic and Statistical Manual of Mental Disorders (DSM)
54
The Diagnostic and Statistical Manual of Mental Disorders (DSM) serves as the primary classification system for mental health disorders, providing standardized diagnostic criteria for clinicians and researchers. First published by the American Psychiatric Association (APA) in 1952, the DSM has undergone several revisions to reflect evolving psychiatric understanding. The fifth edition, DSM-5, released in 2013, introduced key updates that expanded diagnostic categories and modified diagnostic...
54

