Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Data Validation01:03

Data Validation

6.3K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
6.3K
Language Development01:22

Language Development

806
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
806
Language and Cognition01:27

Language and Cognition

693
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
693
Reliability and Validity01:29

Reliability and Validity

13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Components of Language01:24

Components of Language

723
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
723

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Periarticular Embolization as an Alternative Treatment for Surgery-Ineligible Patients with Hip Osteoarthritis: A Prospective Comparative Study.

Journal of clinical medicine·2026
Same author

The Impact of Mixing Techniques on PMMA Bone Cement Subjected to Two Different Cooling Techniques: A Pilot Study of Thermal Management Strategies in Orthopedic Applications.

Biomedicines·2025
Same author

Early Outcomes of Cruciate-Retaining Versus Posterior-Stabilized Total Knee Arthroplasty in Younger Patients: A Prospective Eastern European Cohort Study.

Journal of clinical medicine·2025
Same author

Can PRP Enhance Hamstring Recovery Post-ACL Reconstruction? Retrospective Insights from Non-Professional Athletes.

Journal of clinical medicine·2025
Same author

Single-Session Bilateral Genicular Artery Embolization for Knee Osteoarthritis via Brachial Access: A Case Report and Literature Review.

Diagnostics (Basel, Switzerland)·2025
Same author

Exploring Named Entity Recognition Potential and the Value of Tailored Natural Language Processing Pipelines for Radiology, Pathology, and Progress Notes in Clinical Decision Support: Quantitative Study.

JMIR AI·2025

相关实验视频

Updated: Jan 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

994

临床大语言模型评估通过专家审查 (CLEVER):框架开发和验证.

Veysel Kocaman1, Mustafa Aytuğ Kaya2, Andrei Marian Feier1

  • 1John Snow Labs Inc, 16192 Coastal Highway, Lewes, DE, 19958, United States, +1 (302) 786-5227.

JMIR AI
|December 4, 2025
PubMed
概括

一种新的评估方法CLEVER表明,一个专门的小型LLM在临床任务中表现优于GPT-4o. 这突显了医疗应用中医疗专用大语言模型 (LLM) 的潜力.

关键词:
在NLP中,我们使用了NLP.人工智能的人工智能是人工智能.临床相关性临床相关性评估框架 评估框架事实上的事实性.生成型的人工智能大型语言模型.医学LLM 医学LLM 医学LLM自然语言处理自然语言处理.

更多相关视频

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.3K
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.5K

相关实验视频

Last Updated: Jan 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

994
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.3K
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.5K

科学领域:

  • 人工智能在医学中的应用
  • 自然语言处理自然语言处理.
  • 临床信息学 临床信息学

背景情况:

  • 评估大型语言模型 (LLM) 是一个挑战,因为数据污染和基准任务和临床实践之间的差距.
  • 现有的LLM评估方法,如公共基准和LLM-as-a-judge,受到数据问题和自我偏好的偏见的限制.
  • 迫切需要强大的评估框架,反映现实世界的临床实用性.

研究的目的:

  • 引入CLEVER (临床大语言模型评估-专家评审),这是评估医疗保健LLM的新方法.
  • 使用执业医生对LLM进行盲目,随机,基于偏好的评估.
  • 为了比较一般目的的LLM与医疗保健特定的LLM在临床任务上的表现.

主要方法:

  • 采用CLEVER方法来比较GPT-4o与两个医疗保健特定的LLM (8B和70B参数).
  • 对三个不同的临床任务进行了评估:文本总结,信息提取和回答问题.
  • 执业医生提供基于偏好的评估,重点关注事实性,临床相关性和简洁性.

主要成果:

  • 在关键临床维度中,医生在45%至92%的病例中更喜欢较小的,针对医疗保健的LLM而不是GPT-4o.
  • 与GPT-4o相比,特定于医疗保健的LLM在事实性,临床相关性和简洁性方面表现优异.
  • 在开放式的医疗问题回答方面,表现相似,这表明专业的LLM在上下文依赖的任务中表现出色.

结论:

  • 特定于医疗保健的LLM可以在需要了解临床背景的任务中超过更大,通用目的的LLM.
  • CLEVER方法为评估临床LLM提供了有效和可靠的方法,通过注释者间协议和相关性分析得到证实.
  • 这项研究强调了专门的LLM和专家评审对于推进AI在医学中的重要性.