Consistent performance between medical experts and non-expert readers in forced-choice lesion-detection tasks with
Craig K Abbey1, Sangtae Ahn2, Muhan Shao2
1Department of Psychological and Brain Sciences, University of California, Santa Barbara, California, USA.
Background:
Labeled data are used to train, validate, and test deep learning model observers (DLMOs) as well as linear model observers, such as channelized Hotelling observers (CHOs), for image quality assessment in many imaging modalities, including PET imaging. Ideally, these annotations would come from clinical readers, but these readers are often difficult to access, and it is unclear whether clinical experience is necessary in simple detection tasks involving a known signal profile at a specified location in an image using patient data for backgrounds.
Purpose:
We evaluate the potential for using non-medical observers in simple detection tasks to train DLMOs and CHOs. To this end, we compare nuclear medicine physicians to medically naïve readers using a two-alternative forced choice task with PET images and simulated low-contrast lesions.
Methods:
Experimental conditions span 2 locations in the body (liver and lung), 3 lesion contrasts, 3 implementations of 2 different reconstruction algorithms based on ordered subsets expectation maximization and penalized likelihood. The image readers consist of four board-certified nuclear medicine physicians and four non-physician readers. We make direct comparisons of observer performance between the two reader groups and analyze the observer performance data using generalized linear mixed models. In addition, we train two DLMOs using the expert reader data and the non-expert reader data, respectively, and compare the performance of the DLMOs in predicting the performance of (held-out) expert readers. The DLMOs are based on a CNN and a combination of CNN and Transformer (SwinT) encoders. We also compare CHOs trained on the expert reader data and the non-expert reader data.
Results:
Similarities among the expert readers and between expert and non-expert readers, measured by a concordance metric and Cohen's kappa, were not found to be significantly different (p > 0.11). Analysis of variance found that reader expertise was only marginally significant (p = 0.07) and all interactions involving expertise were non-significant (p > 0.2). Furthermore, the DLMOs trained on experts and on non-experts showed similar performance in predicting experts for both CNN and CNN-SwinT DLMOs (p > 0.26). The CHOs trained on experts and on non-experts showed similar performance (p > 0.31).
Conclusions:
We did not find substantive differences between the average performance of medical experts and non-experts. This suggests that in simple detection tasks like these, clinical experience is not necessary, and we can use non-medical observers instead of physicians when training models.
More Related Videos
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
08:40Positron Emission Tomography Imaging for In Vivo Measuring of Myelin Content in the Lysolecithin Rat Model of Multiple Sclerosis
Published on: February 28, 2021
Related Concept Videos
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body being...
Radiological Investigation III: Pulmonary Angiogram and PET Scan
Pulmonary Angiogram
A Pulmonary Angiogram is an invasive procedure involving injecting a contrast medium through a catheter threaded into the pulmonary artery or the right side of the heart to visualize the pulmonary vasculature. Computed Tomography (CT) scans have mainly replaced this...
Imaging Studies II: Positron Emission Tomography and Scintigraphy
Fundamental Principles of PET
