Related Experiment Video
Updated: Jan 28, 2026

Production of Synthetic Nuclear Melt Glass
Published on: January 4, 2016
High-accuracy prediction of mental health scores from English BERT embeddings trained on LLM-generated synthetic
Birger Moëll1, Fredrik Sand Aronsson2,3
1Division of Speech, Music and Hearing, School of Electrical Engineering and Computer Science, KTH Royal Institute of Technology, Stockholm, Sweden.
Objective:
To assess whether synthetic-only first-person clinical self-reports generated by a large language model (LLM) can support accurate prediction of standardized mental-health scores, enabling a privacy-preserving path for method development and rapid prototyping when real clinical text is unavailable.
Methods:
We prompted an LLM (Gemini 2.5; July 2025 snapshot) to produce English-language first-person narratives that are paired with target scores for three instruments-PHQ-9 (including suicidal ideation), LSAS, and PCL-5. No real patients or clinical notes were used. Narratives and labels were created synthetically and manually screened for coherence and label alignment. Each narrative was embedded using bert-base-uncased (mean-pooled 768-d vectors). We trained linear/regularized linear (Linear, Ridge, Lasso) and ensemble models (Random Forest, Gradient Boosting) for regression, and Logistic Regression/Random Forest for suicidal-ideation classification. Evaluation used 5-fold cross-validation (PHQ-9/SI) and 80/20 held-out splits (LSAS/PCL-5). Metrics: MSE, , MAE; classification metrics are reported for SI.
Results:
Within the synthetic distribution, models fit the label-text signal strongly (e.g., PHQ-9 Ridge: MSE , ; LSAS Gradient Boosting test: MSE , ; PCL-5 Ridge test: MSE , ).
Conclusions:
LLM-generated self-reports encode a score-aligned signal that standard ML models can learn, indicating utility for privacy-preserving, synthetic-only prototyping. This is not a clinical tool: results do not imply generalization to real patient text. We clarify terminology (synthetic text vs. real text) and provide a roadmap for external validation, bias/fidelity assessment, and scope-limited deployment considerations before any clinical use.
More Related Videos
Related Concept Videos
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Stress and Mental Health
Individuals with depression often experience challenges in both their personal and professional...
Opioid Analgesics: Synthetic and Semisynthetic Opioids
Uncertainty in Measurement: Accuracy and Precision
Improving Translational Accuracy
Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT

