Related Experiment Video
Updated: Oct 1, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
An interpretable computational phenotyping pipeline for adult autism using routine electronic health records
Jordi Rodeiro-Boliart1,2,3, Maria Nuñez-Jimenez4, Beatriz Olaya1,5,6
1Impacte i prevenció dels trastorns mentals, Institut de Recerca Sant Joan de Déu, Sant Boi de Llobregat, Spain.
Introduction:
Electronic health records (EHRs) provide new opportunities for computational phenotyping in mental health, but routinely collected diagnostic data are often sparse, heterogeneous, and challenging to interpret. While numerous computational approaches have been applied to identify patient subgroups, clinically useful phenotyping requires workflows capable of transforming routine diagnostic information into interpretable patient profiles. The goal of this work is to develop and evaluate an end-to-end computational phenotyping pipeline for routine ICD-10 diagnostic data and demonstrate its feasibility in adults with autism spectrum disorder (ASD).
Methods:
A retrospective observational study was conducted using routine EHR data from adults with ASD receiving care at a specialized mental health service. ICD-10 diagnoses were extracted, aggregated into clinically meaningful diagnostic categories, and transformed into binary patient-level representations. A self-organizing map (SOM) was used to organize patients according to diagnostic similarity, followed by hierarchical clustering to identify phenotype groups. Cluster solutions were evaluated using internal validity metrics, cluster size distributions, and interpretability. The resulting phenotypes were identified through diagnostic prevalence profiles and SOM-based visualizations.
Results:
The study included 927 adults with ASD, of whom 744 presented at least one comorbid diagnostic category and were included in the computational phenotyping analysis. The proposed pipeline identified four phenotypes characterized by distinct patterns of psychiatric and developmental comorbidity. The SOM structure provided an interpretable visualization of the diagnostic landscape, while hierarchical clustering enabled the identification of coherent phenotype groups. Internal validation metrics supported the selected solution while preserving informative subgroup sizes.
Discussion:
This study presents an end-to-end computational phenotyping pipeline that transforms routinely collected psychiatric EHR data into interpretable patient phenotypes. By integrating diagnostic extraction, clinically informed aggregation, representation learning, clustering, visualization, and characterization, the workflow provides a practical framework for analyzing complex diagnostic data in real-world mental health settings. Future studies should evaluate its applicability across other psychiatric and neurodevelopmental populations.