Fast Model Adaptation for Automated Section Classification in Electronic Medical Records
Jian Ni1, Brian Delaney2, Radu Florian1
1IBM T. J. Watson Research Center, Yorktown Heights, NY, USA.
Studies in Health Technology and Informatics
|August 12, 2015
Summary
This study introduces active learning and distant supervision to reduce the cost and time for training medical information extraction models. These machine learning methods significantly cut annotation expenses for section classification in electronic medical records.
Area of Science:
- Medical Informatics
- Machine Learning
- Natural Language Processing
Background:
- Medical information extraction automates data retrieval from electronic medical records for improved healthcare.
- Section classification is a key task, identifying and categorizing sections within medical documents.
- High accuracy in section classification models requires extensive human-labeled data, which is costly and time-consuming to acquire.
Purpose of the Study:
- To reduce the annotation cost and time for training section classification models.
- To enable faster adaptation of section classification models to new medical institutions.
- To explore the effectiveness of active learning and distant supervision for automated medical information extraction.
Main Methods:
- Applied active learning to intelligently select data for human labeling, minimizing annotation effort.
- Utilized distant supervision to leverage existing, weakly labeled data for model training.
- Evaluated the performance of these techniques on the section classification task in electronic medical records.
Main Results:
- Active learning reduced annotation cost and time by over 50%.
- Distant supervision achieved good model accuracy using only weakly labeled data.
- Demonstrated the feasibility of adapting section classification models efficiently to new institutions.
Conclusions:
- Active learning and distant supervision are effective strategies for reducing annotation burdens in medical information extraction.
- These methods facilitate rapid and cost-effective deployment of section classification models in diverse healthcare settings.
- The study highlights a path towards more scalable and accessible automated medical record analysis.
