Related Experiment Video
Updated: Jun 3, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Open-source Large Language Models can Generate Labels from Radiology Reports for Training Convolutional Neural
Fares Al Mohamad1, Leonhard Donle2, Felix Dorfner3
1Department of Radiology, Charité - Universitätsmedizin Berlin, corporate member of Freie Universität Berlin and Humboldt-Universität zu Berlin, Charitéplatz 1, 10117 Berlin, Germany (F.A.M., L.D., F.D., L.R., K.D., H.H., L.X., F.B.).
Large language models (LLMs) can automatically classify radiology reports, generating labels to train Convolutional Neural Networks (CNNs) for detecting ankle fractures. This automated labeling shows high accuracy, comparable to manual methods, streamlining medical image analysis.
Area of Science:
- Medical Imaging Analysis
- Artificial Intelligence in Healthcare
- Natural Language Processing
Background:
- Training Convolutional Neural Networks (CNNs) for medical image analysis requires extensive labeled datasets, which are time-consuming and costly to create.
- Radiology reports contain valuable information but are often unstructured, hindering direct use in machine learning models.
- Large Language Models (LLMs) offer a potential solution for interpreting and extracting information from unstructured text data.
Purpose of the Study:
- To explore the efficacy of using LLMs to classify radiology reports and generate labels for training medical imaging models.
- To develop and evaluate a CNN trained on LLM-generated labels for the detection of ankle fractures.
- To assess the performance of automated labeling against traditional manual labeling methods.
Main Methods:
- An open-weight LLM (Mixtral-8×7B-Instruct-v0.1) was employed to classify radiology reports of ankle X-ray images.
- Prompt engineering techniques were utilized to optimize the LLM's classification accuracy.
- Labels generated by the LLM were used to train a CNN for ankle fracture detection.
Main Results:
- The LLM achieved 92% accuracy in classifying radiology reports on a test dataset.
- A training dataset of 15,896 images and labels was automatically generated using the LLM.
- The CNN trained on this dataset achieved 89.5% accuracy and an area under the receiver operating characteristics curve of 0.926 for ankle fracture detection.
Conclusions:
- LLM-based automated labeling provides a highly accurate and efficient method for preparing datasets for medical image analysis.
- The performance of models trained with LLM-generated labels is comparable to those trained with manually labeled data.
- LLMs demonstrate significant potential in automating the detection of pathologies within radiology reports, enhancing diagnostic workflows.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020