Related Experiment Video
Updated: Jul 26, 2025

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Supervised Text Classification System Detects Fontan Patients in Electronic Records With Higher Accuracy Than ICD
Yuting Guo1, Mohammed A Al-Garadi2, Wendy M Book3,4
1Department of Biomedical Informatics, School of Medicine Emory University Atlanta GA.
Natural language processing models accurately identify Fontan patients from electronic health records, outperforming traditional International Classification of Diseases (ICD) codes. This advancement aids in creating larger Fontan patient cohorts for research.
Area of Science:
- Medical Informatics
- Computational Biology
- Clinical Data Analysis
Background:
- The Fontan operation is linked to significant patient morbidity and mortality.
- Identifying Fontan patients using International Classification of Diseases (ICD) codes is challenging due to limitations in code accuracy.
- Accurate identification of Fontan cases is crucial for research and clinical management.
Purpose of the Study:
- To develop and compare natural language processing (NLP)-based machine learning models for automatic Fontan patient detection.
- To evaluate the performance of NLP models against traditional ICD code-based classification.
- To enhance the creation of large Fontan patient cohorts for research.
Main Methods:
- Utilized free-text clinical notes from 10,935 manually validated patients (7.1% Fontan).
- Trained and optimized machine learning models, including Support Vector Machines (SVM) and RoBERTa (a transformer-based language model), on 80% of the data.
- Implemented a novel sliding window strategy for RoBERTa to manage text length limitations and evaluated models on held-out data using F1 score.
Main Results:
- Support Vector Machines achieved the highest F1 score (0.95), outperforming both RoBERTa (0.89) and ICD code classification (0.81).
- Both NLP models demonstrated significantly higher accuracy than ICD code-based classification (P<0.05).
- The sliding window strategy improved RoBERTa's performance but did not surpass SVM; ICD codes generated more false positives.
Conclusions:
- NLP models can automatically identify Fontan patients from clinical notes with superior accuracy compared to ICD codes.
- These findings highlight the potential of NLP in improving patient cohort identification for Fontan-associated research.
- Further advancements in NLP offer promising avenues for more precise clinical data analysis.
More Related Videos
13:44Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Heart Failure IV: Classification and Diagnostic Evaluation
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Cardiomyopathy V: Interprofessional Care