Related Experiment Video
Updated: May 3, 2026

Quantitative Optical Microscopy: Measurement of Cellular Biophysical Features with a Standard Optical Microscope
Published on: April 7, 2014
A Method for Efficient De-identification of DICOM Metadata and Burned-in Pixel Text
Jacob A Macdonald1, Katelyn R Morgan2, Brandon Konkel2
1Department of Radiology, Duke University, Durham, NC, USA. jacob.macdonald@duke.edu.
This study introduces an efficient method for de-identifying medical images, significantly reducing the need for time-consuming optical character recognition (OCR) on images containing protected health information (PHI). The approach successfully de-identified over 400,000 DICOM images, enhancing medical research data security.
Area of Science:
- Medical Imaging
- Data Security
- Radiology Informatics
Background:
- De-identification of DICOM images is crucial for medical research.
- Existing methods for removing protected health information (PHI) from DICOM metadata are established, but removing PHI
- burned-in
- to pixel data is often manual and lacks validated high-throughput automation.
- Optical character recognition (OCR) models can detect PHI in images but are computationally intensive for large datasets.
Purpose of the Study:
- To develop and validate an efficient, high-throughput method for de-identifying DICOM images, addressing the challenge of
- burned-in
- PHI.
- To combine automated metadata de-identification with a targeted OCR approach for images likely to contain PHI in pixel data.
Main Methods:
- A novel data processing method was implemented, performing metadata de-identification on all images.
- A targeted strategy was employed to apply OCR exclusively to images identified as having a high probability of containing burned-in PHI.
- The method was validated on a large dataset comprising 415,182 DICOM images across ten modalities.
Main Results:
- The validated method demonstrated high efficacy in de-identifying DICOM images.
- Out of 12,578 images with any burned-in text, only 10 instances were missed by the method.
- OCR was applied to only 6,050 images (1.5% of the total dataset), significantly improving processing efficiency.
Conclusions:
- The developed method provides an efficient and effective solution for de-identifying large volumes of medical images.
- This approach significantly reduces the computational burden associated with OCR for burned-in PHI in medical research.
- The validated method enhances the security and usability of medical image datasets for research purposes.
More Related Videos
09:21Human Brown Adipose Tissue Depots Automatically Segmented by Positron Emission Tomography/Computed Tomography and Registered Magnetic Resonance Images
Published on: February 18, 2015
10:39A Label-Free Segmentation Approach for Intravital Imaging of Mammary Tumor Microenvironment
Published on: May 24, 2022