Related Experiment Video
Updated: Jan 17, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Development and validation of natural language processing algorithms in the national ENACT network
Yanshan Wang1,2,3, Jordan Hilsman1,2, Chenyu Li1,2,3
1Clinical and Translational Science Institute, University of Pittsburgh, Pittsburgh, PA, USA.
The ENACT NLP Working Group successfully deployed natural language processing infrastructure across 13 sites, enabling access to clinical narratives for translational research. This demonstrates the feasibility of federated NLP deployment and highlights the importance of addressing data heterogeneity.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Translational Research
Background:
- Electronic Health Record (EHR) data are crucial for advancing translational research and AI.
- Clinical narratives within EHRs contain valuable information requiring Natural Language Processing (NLP) for research.
- The ENACT network aims to provide access to structured EHR data across 57 Clinical and Translational Science Awards (CTSA) hubs.
Purpose of the Study:
- To establish and operationalize the ENACT NLP Working Group for making NLP-derived clinical information accessible and queryable across the ENACT network.
- To develop and validate NLP algorithms for diverse clinical tasks within specific disease contexts.
- To extend the ENACT ontology and common data model to incorporate NLP-derived data while ensuring compatibility with existing research networks like SHRINE.
Main Methods:
- Established the ENACT NLP Working Group with 13 sites selected based on access to clinical notes, IT infrastructure, NLP expertise, and institutional support.
- Organized sites into five focus groups targeting specific clinical tasks, with each group comprising development and validation sites.
- Extended the ENACT ontology, standardized NLP-derived data, and conducted multisite evaluations using the Open Health Natural Language Processing (OHNLP) Toolkit.
Main Results:
- Achieved 100% site retention and successfully deployed NLP infrastructure across all participating sites.
- Developed and validated NLP algorithms for phenotyping rare diseases, social determinants of health, opioid use disorder, sleep, and delirium.
- Observed performance variability (F1 scores 0.53-0.96) across sites, underscoring the impact of data heterogeneity on NLP model generalizability.
Conclusions:
- Demonstrated the feasibility of deploying NLP infrastructure across large, federated research networks.
- The focus group approach was found to be more practical than general-purpose NLP strategies.
- Key challenges identified include data heterogeneity and the need for robust collaborative governance, providing a foundation for other networks to implement NLP for translational research.
More Related Videos
Related Concept Videos
Automatic Processing and Automatic Social Behavior
Natural and Artificial Concepts
Nonsense-mediated mRNA Decay
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...

