Related Experiment Videos
Towards a Compositional Framework for Describing Human Phenotypes
Wanting Hu1, Yunchao Ling1, Xinyue Hu1
1Bio-Med Big Data Center, Shanghai Institute of Nutrition and Health, University of the Chinese Academy of Sciences, Chinese Academy of Sciences, Shanghai, China.
Abstract:
Deep phenotyping increasingly spans scales from macro to micro, yet heterogeneity in measurement and documentation still constrains standardization and limits comparability across cohorts. This study introduces a compositional framework that integrates the Phenotype Assembly Method (PhenoAM), Phenome Data Elements (PhenoDE), and a suite of large language model (LLM)-based PhenoAgents to support systematic, interpretable, and scalable phenotype standardization with quality assessment. PhenoAM decomposes phenotype descriptions into features and qualifiers to enable component-level mapping, while PhenoDE provides standardized representations of 58 371 phenotype data elements across 22 measurement platforms of the International Human Phenome Project (IHPP); pooled mapping coverage was 98.1% within the 18-platform analysis set and 99.8% across all 22 platforms. Using this representation, analyses of illustrative phenotypes from IHPP, National Health and Nutrition Examination Survey (NHANES), and the UK Biobank show that differences in features or qualifiers can influence data distributions and analytic outcomes, underscoring the importance of explicit measurement context in cross-cohort studies. PhenoAgents support semantic parsing, terminology candidate retrieval, and data-quality review through curator-assisted modules. PhenoDEs and their statistical distributions within IHPP are accessible on the PhenoDE Portal.