Related Experiment Video
Updated: May 8, 2026

Generation of Comprehensive Thoracic Oncology Database - Tool for Translational Research
Published on: January 22, 2011
An Integrated Pipeline for Phenotypic Characterization, Clustering and Visualization of Patient Cohorts in a Rare
Xiaoyi Chen1,2,3, Junyuan Wang1, Carole Faviez2,3,4
1Data Science Platform, Imagine Institute, Université Paris Cité, Inserm UMR 1163, Paris, France.
Abstract:
Rare diseases pose significant challenges due to their heterogeneity and lack of knowledge. This study develops a comprehensive pipeline interoperable with a document-oriented clinical data warehouse, integrating cohort characterization, patient clustering and interpretation. Leveraging NLP, semantic similarity, machine learning and visualization, the pipeline enables the identification of prevalent phenotype patterns and patient stratification. To enhance interpretability, discriminant phenotypes characterizing each cluster are provided. Users can visually test hypotheses by marking patients exhibiting specific keywords in the EHR like genes, drugs and procedures. Implemented through a web interface, the pipeline enables clinicians to navigate through different modules, discover intricate patterns and generate interpretable insights that may advance rare diseases understanding, guide decision-making, and ultimately improve patient outcomes.

