Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Multiple Comparison Tests01:13

Multiple Comparison Tests

4.5K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.5K
Feedback Inhibition00:46

Feedback Inhibition

57.3K
Biochemical reactions are occurring constantly in cells, converting starting substances to different products, usually with the help of enzymes that speed the reactions. Without enzymes, it would take far too long for most reactions to occur to be useful to the cell!
57.3K
Feedback Loops01:01

Feedback Loops

64.7K
In most cases, excessive hormone production is prevented by negative feedback—a loop that starts with a stimulus inducing the release of a particular substance, like a hormone, to maintain a certain level before triggering a signal that results in a decrease in further release of the hormone.
64.7K
Distance Corrections01:15

Distance Corrections

301
To achieve precise distance measurements, especially in surveying and construction, certain corrections must be applied to account for potential sources of error like the standardization errors, temperature variations, and slope adjustments.Standardization error emerges when measurement equipment undergoes changes, such as wear, repairs, or weather impacts. To address this, surveyors compare the equipment’s readings to a standard. This process identifies any deviation that might lead to...
301
Power Factor Correction01:20

Power Factor Correction

557
The power transmission to a factory involves the transfer of apparent power, a combination of active and reactive power. The power factor measures how effectively electrical power is converted into useful work output. The ratio of the real power (KW) that does the work to the apparent power (KVA) supplied to the circuit.
557
Mate Choice01:20

Mate Choice

11.8K
Mate choice—the decision about whom to mate with—is a type of natural selection, since animals must reproduce to pass down their genes. Mate choice is also called intersexual selection because the behavior occurs between the sexes.
11.8K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine.

Findings of ACL. ACL·2026
Same author

RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models.

Proceedings of the conference. Association for Computational Linguistics. Meeting·2026
Same author

MedQA-CS: Objective Structured Clinical Examination (OSCE)-Style Benchmark for Evaluating LLM Clinical Skills.

Proceedings of the conference. Association for Computational Linguistics. European Chapter. Conference·2026
Same author

ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes Care.

Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence·2026
Same author

Social Determinants of Health and 1-Year Buprenorphine Initiation Among Justice-Involved Veterans.

JAMA network open·2026
Same author

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations.

Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing·2026

相关实验视频

Updated: Feb 13, 2026

The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies
08:24

The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies

Published on: August 25, 2023

1.2K

MCQG-SRefine:多选题生成和评估与代自我批评,纠正和比较反.

Zonghai Yao1, Aditya Parashar1, Huixue Zhou2

  • 1University of Massachusetts, Amherst.

Proceedings of the conference. Association for Computational Linguistics. North American Chapter. Meeting
|February 12, 2026
PubMed
概括

本研究介绍了MCQG-SRefine,这是一种使用大型语言模型 (LLM) 来生成高质量的多选择题 (MCQG) 的新型框架,用于医疗执照考试. 该方法改善了问题质量和难度评估.

更多相关视频

A Two-interval Forced-choice Task for Multisensory Comparisons
07:13

A Two-interval Forced-choice Task for Multisensory Comparisons

Published on: November 9, 2018

11.5K
Control of Eating Behavior Using a Novel Feedback System
04:48

Control of Eating Behavior Using a Novel Feedback System

Published on: May 8, 2018

11.7K

相关实验视频

Last Updated: Feb 13, 2026

The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies
08:24

The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies

Published on: August 25, 2023

1.2K
A Two-interval Forced-choice Task for Multisensory Comparisons
07:13

A Two-interval Forced-choice Task for Multisensory Comparisons

Published on: November 9, 2018

11.5K
Control of Eating Behavior Using a Novel Feedback System
04:48

Control of Eating Behavior Using a Novel Feedback System

Published on: May 8, 2018

11.7K

科学领域:

  • 人工智能的人工智能
  • 自然语言处理自然语言处理.
  • 医学教育 医学教育

背景情况:

  • 自动问题生成 (QG) 对人工智能和NLP应用,如智能辅导和对话系统至关重要.
  • 对于专业考试,如美国医学执照考试 (USMLE),生成高质量的多选择题 (MCQG) 是具有挑战性的,因为该领域的专业知识和推理要求.
  • 当前的大型语言模型 (LLM) 在专业的MCQG中面临局限性,包括过时的知识,幻觉和快速敏感性,导致问题质量和难度低于最佳.

研究的目的:

  • 开发一个基于LLM自我精炼的框架 (MCQG-SRefine) 来从医疗案例中生成高质量的USMLE样式问题.
  • 通过专家驱动的快速工程和代的自我批评和自我纠正来提高生成问题的质量和难度.
  • 引入使用LLM-as-Judge的自动评估指标,以取代昂贵的专家评估.

主要方法:

  • 提出了MCQG-SRefine框架,整合了LLM自我批评和纠正机制.
  • 雇佣专家驱动的快速工程来指导问题生成过程.
  • 开发了一个LLM-as-Judge指标,用于自动评估问题质量和难度.

主要成果:

  • MCQG-SRefine显著提高了人类专家对生成USMLE样式问题的质量和难度的满意度.
  • 该框架有效地将医疗病例转化为具有挑战性和相关的MCQ.
  • 作为法官的LLM指标证明了可靠和与专家一致的评估,为手动专家审查提供了替代方案.

结论:

  • MCQG-SRefine为生成高质量的医疗MCQG提供了强大的解决方案,解决了当前LLMs的局限性.
  • 自我改进方法提高了问题相关性和难度,这对于专业考试至关重要.
  • 使用LLM-as-Judge的自动评估提供了一个可扩展和具有成本效益的方法来评估生成的问题.