对诊断准确性研究的偏差风险评估使用大型语言模型的 QUADAS 2 来进行诊断准确性研究
Daniel-Corneliu Leucuța1, Andrada Elena Urda-Cîmpean1, Dan Istrate1
1Department of Medical Informatics and Biostatistics, Iuliu Hațieganu University of Medicine and Pharmacy, 400349 Cluj-Napoca, Romania.
Diagnostics (Basel, Switzerland)
|June 26, 2025
概括
大型语言模型 (LLM) 在使用 QUADAS 2 的诊断准确性研究中,在评估偏差风险方面表现出中度的准确性. 虽然不取代人类专家,但LLM可以帮助监督系统审查.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 生物统计学 生物统计学
背景情况:
- 诊断准确性研究对于评估医疗测试性能至关重要.
- 在这些研究中,偏差风险 (RoB) 评估通常使用诊断准确性研究 (QUADAS) 的质量评估工具.
- 在RoB评估中评估大型语言模型 (LLM) 的功能是一个新的研究领域.
研究的目的:
- 在使用QUADAS 2的诊断准确性研究中评估RoB时评估LLMs的准确性.
- 将LLM绩效与人类专家评估进行比较.
- 确定LLM在RoB评估中表现出色或失败的特定领域.
主要方法:
- 四个LLM (ChatGPT 4o,Grok 3,Gemini 2.0 Flash,DeepSeek V3) 已经被使用.
- 十项诊断准确性研究被选中进行评估.
- 人类专家和法学士独立地将QUADAS 2工具应用于每个研究.
主要成果:
- 在110个信号问题中,LLMs的平均准确率为72.95%.
- 格洛克3 (74.45%) 和ChatGPT 4o (73.15%) 的准确性比DeepSeek V3 (70.00%) 和Gemini 2.0 Flash (67.27%) 的准确性更高.
- 在"流量和时间"领域观察到最高的准确性,其次是"索引测试"",患者选择"和"参考标准"领域,有记录的推理错误.
结论:
- 在诊断准确性研究中,LLM在RoB评估中表现出适度的能力.
- 目前的LLM绩效不足以取代专家的临床和方法论判断.
- 在系统性审查中,LLM显示出作为补充工具的潜力,这取决于强制性的人类监督.
相关概念视频
Detection of Gross Error: The Q Test
6.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.4K
Bias in Epidemiological Studies
705
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
705
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Sensitivity, Specificity, and Predicted Value
683
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
683
Accuracy and Errors in Hypothesis Testing
322
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
322
Receiver Operating Characteristic Plot
339
A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
339


