Related Experiment Video
Updated: Jul 4, 2025

12:32
Image Rendering Techniques in Postmortem Computed Tomography: Evaluation of Biological Health and Profile in Stranded Cetaceans
Published on: September 27, 2020
8.7K
Inconsistency between Human Observation and Deep Learning Models: Assessing Validity of Postmortem Computed
Yuwen Zeng1, Xiaoyong Zhang2, Jiaoyang Wang3
1Department of Radiological Imaging and Informatics, Tohoku University Graduate School of Medicine, Sendai, Japan. yuwen@tohoku.ac.jp.
Journal of Imaging Informatics in Medicine
|February 9, 2024
Summary
Deep learning models for drowning diagnosis show high performance but lack medical validity. Saliency maps reveal irrelevant features, questioning the reliability of these artificial intelligence tools in forensic pathology.
Area of Science:
- Forensic Pathology
- Medical Imaging
- Artificial Intelligence
Background:
- Drowning diagnosis in autopsy is challenging, even with advanced imaging.
- Deep learning (DL) models have shown promise for drowning diagnosis.
- The medical validity of these DL models has not been previously assessed.
Purpose of the Study:
- To assess the medical validity of DL models used for drowning diagnosis.
- To determine if DL models' learned features align with expert radiological findings.
- To evaluate the reliability of high-performing DL tools in forensic autopsy.
Main Methods:
- Retrospective study of autopsy cases (2012-2021) with postmortem computed tomography.
- Trained three DL models and generated saliency maps to highlight important image features.
- Compared DL saliency maps with pixel-level annotations from radiological technologists.
Main Results:
- All three DL models achieved high classification performance (AUCs 0.94-0.98).
- Significant inconsistency was found between model saliency maps and expert annotations.
- Models exhibited 30-80% irrelevant areas in saliency maps, indicating potential unreliability.
Conclusions:
- Despite high classification accuracy, DL models for drowning diagnosis may be unreliable.
- Careful assessment of DL tools is crucial, even when performance metrics are high.
- Findings highlight the need for rigorous validation of AI in forensic medicine.

