Related Experiment Video
Updated: Oct 27, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Structuring clinical text with AI: Old versus new natural language processing techniques evaluated on eight common
Xianghao Zhan1, Marie Humbert-Droz2, Pritam Mukherjee2
1Department of Bioengineering, Stanford University, Stanford, CA 94305, USA.
This study extracts diagnostic codes from clinical notes using word vectorization, improving data mining. TF-IDF vectorization achieved high accuracy, demonstrating feasibility for code extraction and correction.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Machine Learning
Background:
- Electronic health records (EHRs) contain valuable clinical information in free-text notes.
- Structured diagnostic codes in EHRs can be incomplete or inaccurate, hindering data analysis.
- Accurate diagnostic coding is crucial for clinical research, billing, and patient care.
Purpose of the Study:
- To develop and evaluate methods for extracting International Classification of Diseases, 10th Revision (ICD-10) codes from free-text clinical notes.
- To assess the performance and transferability of different word vectorization techniques for diagnostic code prediction.
- To demonstrate the feasibility of improving diagnostic code quality through automated extraction from clinical narratives.
Main Methods:
- Five word vectorization methods (including TF-IDF) were applied to Stanford progress notes.
- Logistic regression was used to predict eight common cardiovascular disease ICD-10 codes.
- Model performance was evaluated using Area Under the Receiver Operating Characteristic Curve (AUROC) and Area Under the Precision-Recall Curve (AUPRC).
- Model transferability was tested on the MIMIC-III database.
Main Results:
- The TF-IDF vectorization model achieved the highest performance, with AUROC ranging from 0.9499 to 0.9915 and AUPRC from 0.2956 to 0.8072.
- Models demonstrated good transferability to MIMIC-III data, with AUROC from 0.7952 to 0.9790 and AUPRC from 0.2353 to 0.8084.
- Analysis of important words provided clinical interpretability for disease prediction.
Conclusions:
- Automated extraction of diagnostic codes from free-text clinical notes is feasible and accurate.
- This approach can help impute missing codes and correct erroneous ones, enhancing EHR data quality.
- The findings support the use of NLP and machine learning for improved information retrieval and downstream applications in healthcare.
More Related Videos
07:51Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Cardiovascular Drugs: Classification based on Therapeutic Indications
Assessment of the Cardiovascular System I: Subjective Data
Initial Enquiry
Ask the patient about their primary concern and thoroughly explore all reported symptoms.
Medical History
Investigate past illnesses affecting the cardiovascular system, such as angina, anemia, rheumatic fever, congenital heart disease, stroke, thrombophlebitis, dysrhythmias, varicosities
Inquire about symptoms...
Pre-Procedural Guidelines for Assessing Blood Pressure
Ischemic Heart Disease: Overview
Atherosclerosis, the primary malefactor, orchestrates this dangerous condition. It manifests as the accumulation of fatty deposits, akin to insidious plaques, within arterial walls. As time elapses, these plaques metamorphose, hardening and...
Overview of the Vascular System
Heart Failure Drugs: β-Blockers