Related Experiment Video
Updated: Jun 21, 2025

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Building RadiologyNET: an unsupervised approach to annotating a large-scale multimodal medical database
Mateja Napravnik1, Franko Hržić1,2, Sebastian Tschauner3
1Faculty of Engineering, University of Rijeka, Vukovarska 58, Rijeka, 51000, Croatia.
This study introduces an automated method to annotate large medical radiology image datasets by analyzing semantic similarity. This approach overcomes the limitations of manual annotation, paving the way for more extensive and detailed medical imaging databases.
Area of Science:
- Medical Imaging
- Machine Learning
- Data Science
Background:
- Machine learning (ML) is increasingly used in medical diagnosis and treatment via computer-aided diagnosis (CAD) systems.
- Annotated medical radiology images are crucial for CAD systems, but manual annotation is time-consuming and expensive.
- A scarcity of large annotated medical image datasets hinders ML development in healthcare.
Purpose of the Study:
- To develop an automated method for annotating large medical radiology image datasets.
- To address the challenge of time-consuming and costly manual image annotation.
- To leverage semantic similarity for efficient medical image dataset annotation.
Main Methods:
- An unsupervised, automated pipeline was developed to create annotated medical radiology image datasets.
- The pipeline integrated data mining of medical images, DICOM metadata, and narrative diagnoses.
- Optimal feature extractors were combined into a multimodal representation, clustered into 50 groups of visually similar images from 1,337,926 images.
Main Results:
- The automated pipeline successfully clustered 1,337,926 medical images into 50 distinct groups based on visual similarity.
- Cluster quality was evaluated using homogeneity and mutual information metrics, considering anatomical region and modality.
- Fusing embeddings from images, DICOM metadata, and narrative diagnoses yielded the most concise and effective clusters.
Conclusions:
- Fusing multimodal data embeddings (images, metadata, diagnoses) is optimal for unsupervised clustering of large-scale medical data.
- This automated approach significantly enhances the efficiency of creating annotated medical image datasets.
- The study represents a foundational step towards developing larger, more granular annotated medical radiology image databases.
More Related Videos
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018