Related Experiment Video
Updated: Sep 12, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
PheCatcher: Leveraging LLM-Generated Synthetic Data for Automated Phenotype Definition Extraction from Biomedical
1McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX, USA.
Abstract:
Phenotype definitions are crucial for the progression of precision and personalized medicine. Although phenotype knowledge bases such as PheKB and the OHDSI library are available, they rely heavily on manual input. This study introduces PheCatcher, an automated pipeline that integrates BiomedBERT-based Named Entity Recognition (NER) and Relation Extraction (RE) to extract phenotypes and standardized codes from biomedical literature. To complement human annotation, GPT-4 was utilized to generate synthetic data, which improved model performance. The NER model's F1 score for "phenotype" entities increased from 0.616 to 0.800, and the RE model achieved an F1 score of 0.901. The application of the pipeline to the PubMed Central (PMC) articles resulted in the extraction of 173,283 phenotype definitions, which are now publicly accessible. Our study underscores the potential of synthetic data for information extraction (IE) and offers the first evidence of the feasibility of leveraging synthetic data to build a complete IE system.
Related Concept Videos
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...

