医学的な質問に対する過剰な自信:定量的な研究
Raphaël Bentegeac1,2, Bastien Le Guellec3,4, Grégory Kuchcinski3,4
1Department of Public Health, Lille University, Lille University Hospital Center, avenue du Professeur Emile Laine, Lille, 59037, France.
Journal of medical Internet research
|August 29, 2025
まとめ
医学的チャットボットは 高い信頼を示しますが 信頼性はありません 自己評価の確実性ではなく トークンによる確率で 医療上の質問のチャットボットの 正確さを予測し より信頼性の高い評価が得られます
科学分野:
- 医療における人工知能
- 自然言語処理
- 医療情報工学
背景:
- チャットボットは医学界で有望で 医学委員会試験に合格しています
- しかし 誤った答えに 過剰な自信があるため 臨床使用は制限されています
研究 の 目的:
- 医学的反応の精度を予測するチャットボットによる信頼を比較する
- チャットボットのパフォーマンスを評価する際に,トークンの確率が自己報告の信頼に優れているかどうかを評価する.
主な方法:
- 9つの大きな言語モデル (LLM) は,米国医療免許試験の2522の質問に答えました.
- モデルの信頼性と応答トークンの確率を記録し分析した.
- 予測性能は,AUROC,校正エラー,およびBrierスコアを使用して評価されました.
主要な成果:
- チャットボットの精度は,GPT-4oが89%,Phi-3-Miniが56.5%を達成した.
- 信頼性の低い予測精度 (AUROC 0.52-0.68) を表した.
- トークン確率は一貫して信頼性 (AUROC 0.71-0.87) を上回り,より優れた精度予測を示しています.
結論:
- チャットボットは医療の文脈で 自信の正確な自己評価に苦労します
- トークン確率はチャットボットの応答の精度を評価するためのより信頼性の高い方法を提供します.
- 臨床医はチャットボットによる自己評価の確実性には 頼るべきではないのです
さらに関連する動画
関連する概念動画
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
The Availability Heuristic
6.4K
A heuristic is a general problem-solving framework (Tversky & Kahneman, 1974). You can think of these as mental shortcuts that are used to solve problems. Different types of heuristics are used in different types of situations, and the impulse to use a heuristic occurs when one of five conditions is met (Pratkanis, 1989):
6.4K
Sensitivity, Specificity, and Predicted Value
658
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
658
Uncertainty: Confidence Intervals
4.6K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
4.6K
Microorganisms in Medicine and Therapeutics
312
Microorganisms play a fundamental role in vaccine development, gene therapy, and therapeutic production. Their biological properties are harnessed to advance medicine and public health. Beyond immunization, microorganisms contribute to gut health, antibiotic synthesis, and genetic disease treatment.Live Attenuated and Inactivated VaccinesLive attenuated vaccines, such as the measles, mumps, and rubella (MMR) vaccine, utilize weakened forms of pathogens to closely resemble natural infections.
312
Testing a Claim about Population Proportion
3.4K
A complete procedure for testing a claim about a population proportion is provided here.
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
3.4K


