医療アプリケーションにおける入力変数による大型言語モデルの性能:データセット開発と実験的評価
Saubhagya Joshi1, Monjil Mehta2, Sarjak Maniar2
1Library and Information Sciences, School of Communication & Information, Rutgers University, 4 Huntington St, New Brunswick, NJ, 08901, United States, +1 (848) 932-7500.
JMIR AI
|February 20, 2026
まとめ
医療における大型言語モデル (LLM) は,タイポスやホモフォンのような一般的な入力エラーに対して驚くほど堅牢であることを示しています. しかし,編集はLLMのパフォーマンスを著しく低下させ,臨床アプリケーションの慎重な設計の必要性を強調します.
科学分野:
- 医療における人工知能
- 自然言語処理 (Natural Language Processing) とは,自然言語処理で処理される言語のことです.
- クリニカル・インフォマティックス
背景:
- 大型言語モデル (LLM) は,医療において,患者のケアと意思決定のためにますます使用されています.
- 不完全な臨床データを持つLLMの信頼性は十分に理解されていません.
- データの不完全性は,臨床文書および患者によって生成された情報において一般的です.
研究 の 目的:
- 医療アプリケーションにおけるLLMパフォーマンスに対する入力混乱の影響を調査する.
- 異なる種類の波動とレベルの波動の影響を比較する.
- 健康関連と健康関連でない用語に対する影響の違いを分析する.
主な方法:
- 3つの健康関連タスクにおける3つのLLMの体系的な評価.
- 編集,ホモフォーン,タイポグラフィカルエラーなど,人間のような変異を伴う新しいデータセットを利用した.
- 各種の perturbation レベルでの評価された性能.
主要な成果:
- LLMは,一般的な入力変数に対する顕著な強度を示し,パフォーマンスは55%以上のケースで安定または改善しました.
- 低レベルの干渉は,時にはパフォーマンスの向上 (14.07%) をもたらしました.
- 編集は,他のバリエーションよりもLLMのパフォーマンスに有害であることが判明しました.
結論:
- LLMを使用するヘルスケアアプリケーションは,入力変動性とデータ品質を考慮する必要があります.
- 不完全な入力に対する堅固さは,臨床環境におけるLLMの信頼性にとって極めて重要です.
- 発見は,回復力のあるAIツールを開発し,医療におけるLLMのパフォーマンスを改善するための洞察を提供します.
関連する概念動画
Improving Translational Accuracy
15.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.2K
Improving Translational Accuracy
3.7K
3.7K
Variability: Analysis
547
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
547
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
277
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
277


