MCQG-SRefine: 繰り返し自己批判,訂正,比較フィードバックによる複数選択質問生成と評価
Zonghai Yao1, Aditya Parashar1, Huixue Zhou2
1University of Massachusetts, Amherst.
まとめ
この研究は,医療免許試験のための高品質の複数選択質問 (MCQG) を生成するために,大型言語モデル (LLM) を使用する新しいフレームワークであるMCQG-SRefineを導入します. この方法は,質問の質と難易度の評価を改善します.
科学分野:
- 人工知能 (AI) とは,人工知能 (AI) のことです.
- 自然言語処理 (Natural Language Processing) とは,自然言語処理で処理される言語のことです.
- 医療教育 医療教育について
背景:
- 自動質問生成 (QG) は,インテリジェント・チュートリング・ダイアログ・システムなどのAIとNLPのアプリケーションにとって極めて重要です.
- 米国医療免許試験 (USMLE) などの専門試験のための高品質の複数の選択質問 (MCQG) を生成することは,ドメインの専門知識と推論要件のために困難です.
- 現在の大型言語モデル (LLM) は,時代遅れの知識,幻覚,迅速な感受性など,プロフェッショナルなMCQGの限界に直面しており,これは最適な質と難易度の低い質問につながります.
研究 の 目的:
- 医療ケースから高品質のUSMLEスタイルの質問を生成するためのLLM自己精錬ベースのフレームワーク (MCQG-SRefine) を開発する.
- 専門家主導のプロンプトエンジニアリングと反復的な自己批判と自己訂正を通じて,生成された質問の質と難易度を向上させる.
- LLM-as-Judgeを用いた自動評価メトリックを導入し,高額な専門家の評価を代用する.
主な方法:
- LLMの自己批判と訂正メカニズムを統合したMCQG-SRefineの枠組みを提案した.
- 質問生成プロセスを導くために,専門家主導のプロンプトエンジニアリングを採用しました.
- 質問の質と難易度の自動評価のためのLLM-as-Judgeメトリックを開発しました.
主要な成果:
- MCQG-SRefineは,生成されたUSMLEスタイルの質問の質と難易度に対する人間の専門家満足度を大幅に改善しました.
- このフレームワークは,医療症例を,挑戦的で関連性のあるMCQに効果的に変換します.
- LLM-as-Judgeメトリックは,信頼性があり,専門家に合わせた評価を示し,手動の専門家のレビューの代替案を提供しました.
結論:
- MCQG-SRefineは,高品質の医療MCQGを生成するための堅牢なソリューションを提供し,現在のLLMsの限界に対処しています.
- 自己精錬のアプローチは,プロフェッショナル試験に不可欠な質問の関連性と難易度を高めます.
- LLM-as-Judgeを使用した自動評価は,生成された質問を評価するためのスケーラブルで費用対効果の高い方法を提供します.
関連する概念動画
Multiple Comparison Tests
4.5K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.5K
Feedback Inhibition
57.3K
Biochemical reactions are occurring constantly in cells, converting starting substances to different products, usually with the help of enzymes that speed the reactions. Without enzymes, it would take far too long for most reactions to occur to be useful to the cell!
57.3K
Feedback Loops
64.7K
In most cases, excessive hormone production is prevented by negative feedback—a loop that starts with a stimulus inducing the release of a particular substance, like a hormone, to maintain a certain level before triggering a signal that results in a decrease in further release of the hormone.
64.7K
Distance Corrections
301
To achieve precise distance measurements, especially in surveying and construction, certain corrections must be applied to account for potential sources of error like the standardization errors, temperature variations, and slope adjustments.Standardization error emerges when measurement equipment undergoes changes, such as wear, repairs, or weather impacts. To address this, surveyors compare the equipment’s readings to a standard. This process identifies any deviation that might lead to...
301
Power Factor Correction
557
The power transmission to a factory involves the transfer of apparent power, a combination of active and reactive power. The power factor measures how effectively electrical power is converted into useful work output. The ratio of the real power (KW) that does the work to the apparent power (KVA) supplied to the circuit.
557
Mate Choice
11.8K
Mate choice—the decision about whom to mate with—is a type of natural selection, since animals must reproduce to pass down their genes. Mate choice is also called intersexual selection because the behavior occurs between the sexes.
11.8K


