Harmonizing Multi-Institutional Clinical Documentation Using Natural Language Processing in Neurofibromatosis Type 1
Stephanie M Morris1,2, Levi Kaster3, Saki Amagai4
1Department of Neurology, Kennedy Krieger Institute, Baltimore, MD.
Neurology. Clinical Practice
|August 12, 2026
Summary
Physician documentation of neurofibromatosis type 1 (NF1) features in electronic health records (EHRs) shows significant variation in terms used and completeness. Developing a standardized clinical lexicon can improve data accuracy for machine learning and clinical trials.
Area of Science:
- Computational linguistics
- Medical informatics
- Genetics and genomics
Background:
- Electronic health records (EHRs) are crucial for machine learning (ML) and natural language processing (NLP) in neurologic disease research.
- Inconsistent clinical documentation, especially in heterogeneous disorders like neurofibromatosis type 1 (NF1), hinders data harmonization and model performance.
- Physician documentation of NF1 features presents systematic lexical variations, impacting computational phenotyping accuracy.
Purpose of the Study:
- To characterize lexical variation and documentation completeness for core NF1 features within EHRs.
- To develop a standardized, data-informed clinical lexicon aligned with current clinical practice and terminology standards.
- To address challenges in NLP and ML-based phenotyping for NF1.
Main Methods:
- Retrospective observational study of outpatient progress notes from pediatric NF1 patients at two large tertiary care centers.
- Development of a rule-based NLP algorithm to identify 10 core NF1 features and extract documented terms.
- Quantification of lexical variants and documentation frequency across institutions, providers, and time.
Main Results:
- Analysis of 5,393 notes from 1,661 pediatric patients revealed substantial lexical variation for most NF1 features.
- Nonstandard terms were frequently used for significant features like optic pathway glioma.
- Documentation completeness varied, with longitudinal data gaps observed over time.
Conclusions:
- Significant lexical variation and incomplete longitudinal capture in physician-authored EHRs limit NLP and ML phenotyping accuracy for NF1.
- A standardized, data-informed clinical lexicon is essential for improving interoperability and phenotypic consistency.
- This strategy can enhance readiness for clinical trials and real-world evidence generation in NF1 and similar complex neurologic disorders.


