Related Experiment Video
Updated: Sep 2, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Multiple Pretrained BERT Models to Automate and Accelerate Data Annotation for Large Datasets.
Ali S Tejani1, Yee S Ng1, Yin Xi1
1Department of Radiology, University of Texas Southwestern Medical Center, 5323 Harry Hines Blvd, Dallas, TX 75390.
Transformer models like BERT can accurately annotate large medical datasets using small training sets and minimal time. This transfer learning approach enhances efficiency in medical informatics and named entity recognition.
Area of Science:
- Medical informatics
- Natural Language Processing
- Machine Learning
Background:
- Automated analysis of clinical text is crucial for extracting valuable information.
- Large datasets require efficient annotation methods to facilitate research and clinical applications.
Purpose of the Study:
- To evaluate domain-specific and pretrained transformer models for medical text annotation.
- To assess the impact of varying training dataset sizes on model performance in a transfer learning context.
Main Methods:
- Retrospective review of 69,095 adult chest radiograph reports.
- Manual annotation of 1004 reports for four device types: endotracheal tube (ETT), nasogastric tube (NGT), central venous catheter (CVC), and Swan-Ganz catheter (SGC).
- Training and validation of multiple transformer models (BERT, PubMedBERT, DistilBERT, RoBERTa, DeBERTa) using varying percentages of the annotated data (5%-40%) and fivefold cross-validation.
Main Results:
- High Area Under the Curve (AUC) values achieved: 0.996 for ETT (RoBERTa), 0.994 for NGT (RoBERTa), 0.991 for CVC (PubMedBERT), and 0.98 for SGC (PubMedBERT).
- DeBERTa model showed the highest AUC when trained on only 5% of the dataset.
- DistilBERT model exhibited the shortest training and validation time (3 minutes 39 seconds).
Conclusions:
- Pretrained and domain-specific transformer models are effective for autonomous annotation of large datasets.
- These models require small training datasets and short training times to achieve high accuracy.
- This approach significantly expedites the annotation process in medical informatics.
More Related Videos
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025