Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Complementation Tests00:49

Complementation Tests

4.9K
A complementation test is a simple cross to identify whether the two mutations are located on the same gene or different genes. It was first performed by Edward Lewis in the 1940s while working on fruit flies. He developed the test to identify the location and arrangement of different mutations on chromosomes.
Organisms heterozygous for different mutations are crossed pairwise in all combinations. If present on different genes, the mutations can complement each other by providing the missing...
4.9K
Multiple Comparison Tests01:13

Multiple Comparison Tests

3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Improving Translational Accuracy02:07

Improving Translational Accuracy

9.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.4K
Sign Test for Nominal Data01:12

Sign Test for Nominal Data

78
The sign test is a nonparametric method used to evaluate hypotheses about the median of a single sample or to compare the medians of two related samples. The sign test is particularly useful when dealing with nominal data, which includes distinct categories without an inherent order, such as names, labels, and preferences. Nominal data restricts statistical analysis to evaluating population proportions rather than mean or median values that require continuous data.
For example, consider a...
78
Sign Test for Matched Pairs01:17

Sign Test for Matched Pairs

117
The sign test for matched pairs offers a robust method for comparing two paired samples, often for the effects of an intervention in one of them. This method is very useful in situations where the underlying distribution of the data is unknown. The test compares two related samples—often pre- and post-treatment measurements on the same subjects—to determine if there are significant differences in their median values.
To conduct the sign test, we first calculate the differences in...
117
Stereotype Content Model02:16

Stereotype Content Model

14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Brain age gradients as intermediate phenotypes linking plasma p-tau217 to cognition in community-dwelling older adults.

NPJ dementia·2026
Same author

Physical function impacts hearing without mediation from systolic blood pressure.

Scientific reports·2026
Same author

Guiding approaches to studying Alzheimer's disease: a scoping review of community engagement, health communication, and implementation science research.

The Gerontologist·2026
Same author

Hypertension-Induced Retinal Microvascular Remodeling in Women.

Journal of vascular research·2026
Same author

Digital Twin Model of Treatment Outcomes in Post-Stroke Aphasia.

medRxiv : the preprint server for health sciences·2026
Same author

Multifaceted neural representation of words in naturalistic language.

ArXiv·2026

相关实验视频

Updated: Jun 13, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

519

两词测试作为大型语言模型的语义基准.

Nicholas Riccardi1, Xuan Yang2, Rutvik H Desai3

  • 1Department of Communication Sciences and Disorders, University of South Carolina, Columbia, 29208, USA.

Scientific reports
|September 16, 2024
PubMed
概括

两个单词测试 (TWT) 揭示了大型语言模型 (LLM) 与基本的语义理解斗争,与人类相比,在含义判断方面表现不佳. 这一基准强调了LLM在真正语言理解方面的局限性.

科学领域:

  • 自然语言处理自然语言处理.
  • 人工智能的人工智能
  • 认知科学 认知科学

背景情况:

  • 大型语言模型 (LLM) 展示了先进的能力,促使人们讨论它们对类似人类理解和人工通用智能 (AGI) 的潜力.
  • 当前的基准通常集中在推理或领域专业知识上,可能会忽视基本的语义处理能力.
  • 人类语言严重依赖于将单词结合起来形成有意义的概念,这是语言的核心操作.

研究的目的:

  • 引入开源的两个词测试 (TWT) 作为一个新的基准来评估LLMs的语义能力.
  • 评估LLM对两个单词短语的有意义判断的能力,这是人类很容易完成的任务.
  • 提供一个工具来识别和解决LLM语言理解的局限性.

主要方法:

  • 开发了两个单词测试 (TWT),包括1768个名词-名词组合,被人类参与者评价为有意义的.
  • 管理TWT的两个版本:一个0-4级别的细微评级和二进制判断任务.
  • 在TWT上测试了包括GPT-4,GPT-3.5,Claude-3-Optus和Gemini-1.0-Pro-001在内的领先的LLM.

主要成果:

  • 所有测试的LLM在判断两个单词短语的意义上表现明显比人类差.
  • GPT-3.5-turbo,Gemini-1.0-Pro-001和GPT-4-turbo无法可靠地区分有意义的和没有意义的短语.

更多相关视频

Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
12:49

Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition

Published on: July 13, 2019

16.8K
Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.1K

相关实验视频

Last Updated: Jun 13, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

519
Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
12:49

Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition

Published on: July 13, 2019

16.8K
Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.1K
  • 克劳德-3-Opus在二元区分方面有所改善,但仍然落后于人类的表现.
  • 结论:

    • 该TWT有效地突出了当前LLMs在基本语义理解方面的局限性.
    • 结果表明,在根据现有基准对LLM赋予人类层面或"真实"理解时,需要谨慎.
    • TWT为评估和潜在地增强未来LLM的语义能力提供了有价值的工具.