Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches01:23

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches

407
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
407

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Anti-Malignancy and Anti-Inflammation Properties of Emodin, an Essential Constituent in Popular Laxative Botanicals.

The American journal of Chinese medicine·2026
Same author

Artificial Intelligence-Enabled Cardiac Function Estimation from Phone Videos of Echocardiograms.

medRxiv : the preprint server for health sciences·2026
Same author

Interpretable spatial multi-omics data integration and dimensionality reduction with SpaMV.

Nature communications·2026
Same author

Reporting Interest-Holder Engagement in Practice Guidelines: The RIGHT-MuSE Checklist.

Annals of internal medicine·2026
Same author

Amino acid-based biological age clock and its implications for human health and aging.

Nature communications·2026
Same author

Integrating Planetary Health in Health Guidelines (GRADE Guidance 46).

Annals of internal medicine·2026

相关实验视频

Updated: Jan 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.0K

使用大型语言模型来评估AI干预随机对照试验与CONSORT-AI的一致性:跨部门调查

Xufei Luo1,2,3,4,5, Zeming Li6, Zhenhua Yang7

  • 1Evidence-Based Medicine Center, School of Basic Medical Sciences, Lanzhou University, 199 Donggang West Road, Chengguan District, Lanzhou, 730000, China, 86 13893104140.

Journal of medical Internet research
|September 26, 2025
PubMed
概括

大型语言模型 (LLM) 在评估研究一致性方面表现有前途. 在评估人工智能 (AI) 随机对照试验 (RCT) 与CONSORT-AI标准相比,GPT-4变体的表现最好,尽管人类监督仍然至关重要.

关键词:
这就是 CONSORT-AI.聊天GPT 聊天GPT 聊天人工智能的人工智能是人工智能.大型语言模型随机对照试验是随机对照试验.

更多相关视频

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.3K

相关实验视频

Last Updated: Jan 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.0K
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.3K

科学领域:

  • 医疗研究中的人工智能
  • 临床试验报告标准 临床试验报告标准
  • 自然语言处理应用程序

背景情况:

  • 大型语言模型 (LLM) 显示了评估研究一致性的潜力.
  • 之前的研究使用了LLM来评估随机对照试验 (RCT) 摘要与CONSORT-Abstract指南相比.
  • 当通过LLM进行评估时,AI干预性RCT遵循CONSORT-AI标准的一致性尚未得到充分确立.

研究的目的:

  • 使用基于LLM的聊天机器人评估AI干预性RCT与CONSORT-AI标准的一致性.
  • 确定不同LLM模型在评估遵守CONSORT-AI指南时的表现.

主要方法:

  • 一项涉及6个LLM模型的横截面研究,以评估JAMA网络开放的41个人工智能干预的RCT.
  • 查询通过API提交,确定性响应的温度设置为0.
  • 分析了整体一致性得分 (OCS),回忆,评审者之间的可靠性和内容一致性,并对LLM响应进行了独立验证.

主要成果:

  • GPT-4变种呈现出最高的平均OCS,gpt-4-0125预览达到86.5% (JAMA作者) 和81.6% (研究作者).
  • GPT-3.5-turbo-0125显示了最低的平均OCS (61.9%和63.0%).
  • CONSORT-AI项目2 ("在输入数据层面列出包含和排除标准") 得到了最差的评估 (48.8%的OCS),而项目1,5,8和9超过了80%的OCS.

结论:

  • GPT-4变种在评估RCT与CONSORT-AI的一致性方面表现出强大的能力.
  • 快速改进是必要的,以提高基于LLM的评估的精度和一致性.
  • 人类监督和专业知识对于医疗研究中可靠的AI驱动评估至关重要,提高了评估效率和质量.