Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Types of Hypothesis Testing01:11

Types of Hypothesis Testing

25.8K
There are three types of hypothesis tests: right-tailed, left-tailed, and two-tailed.
When the null and alternative hypotheses are stated, it is observed that the null hypothesis is a neutral statement against which the alternative hypothesis is tested. The alternative hypothesis is a claim that instead has a certain direction. If the null hypothesis claims that p = 0.5, the alternative hypothesis would be an opposing statement to this and can be put either p > 0.5, p < 0.5, or p...
25.8K
Statistical Hypothesis Testing01:16

Statistical Hypothesis Testing

1.8K
Hypothesis testing is a critical statistical procedure facilitating informed, evidence-based decisions. It begins with a hypothesis, which is a tentative explanation, or a prediction about a population parameter. This hypothesis can be either a null hypothesis (H0), indicating no effect or difference, or an alternative hypothesis (Ha), suggesting an effect or difference.
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
1.8K
Stereotypes, Prejudice, and Discrimination02:55

Stereotypes, Prejudice, and Discrimination

89.7K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
89.7K
Hypothesis: Accept or Fail to Reject?01:17

Hypothesis: Accept or Fail to Reject?

27.4K
The outcome of any hypothesis testing leads to rejecting or not rejecting the null hypothesis. This decision is taken based on the analysis of the data, an appropriate test statistic, an appropriate confidence level, the critical values, and P-values. However, when the evidence suggests that the null hypothesis cannot be rejected, is it right to say, 'Accept' the null hypothesis?
There are two ways to indicate that the null hypothesis is not rejected. 'Accept' the null...
27.4K
Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

157
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
157
Confirmation Biases01:31

Confirmation Biases

5.4K
The confirmation bias is the tendency to focus on information that confirms our existing beliefs and ignore information that is inconsistent with our expectations. For example, if you think that your professor is not very nice, you notice all of the instances of rude behavior exhibited by the professor while ignoring the countless pleasant interactions he is involved in on a daily basis. Have you ever fallen prey to the confirmation bias, either as the source or target of such bias?
5.4K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Haematococcus pluvialis peptides ameliorated cyclophosphamide-induced immunodeficiency in mice by regulating intestinal barrier function.

Journal of the science of food and agriculture·2026
Same author

p38 MAP kinase senses short-chain fatty acids to attenuate Toll-like receptor signaling and intestinal inflammation.

Science advances·2026
Same author

Voice as relational orientation cue: neural dynamics underlying the decoding of listener- and content-oriented vocal communication.

Brain and language·2026
Same author

Value crucible for evaluating robustness of value attributed LLM response profiles via agent adversarial debates.

Scientific reports·2026
Same author

Lactate binds and inhibits the innate immune sensor STING to promote tumor immune evasion.

Immunity·2026
Same author

Molecular Mechanism of Rice Protein Amyloid Fibrils in Modulating Gel Properties of Northern Pike (<i>Esox lucius</i>) Muscle Protein.

Foods (Basel, Switzerland)·2026

相关实验视频

Updated: May 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

465

用贝叶斯假设测试来检测大型语言模型的隐性偏差.

Shijing Si1, Xiaoming Jiang2,3, Qinliang Su4,5

  • 1School of Economics and Finance, Shanghai International Studies University, Shanghai, 201620, China.

Scientific reports
|April 11, 2025
PubMed
概括

本研究引入了一种新的假设测试框架,用于检测大型语言模型 (LLM) 中的社会偏见. 贝叶斯因子有效量化偏见证据,优于传统的统计测试.

关键词:
贝叶斯因子是一个贝叶斯因子.公平的 公平的 公平的群体偏见 群体偏见 群体偏见大型语言模型.

更多相关视频

Post-Movie Subliminal Measurement PMSM, for Investigating Implicit Social Bias
09:03

Post-Movie Subliminal Measurement PMSM, for Investigating Implicit Social Bias

Published on: February 29, 2020

5.7K
Defining the Role Of Language in Infants' Object Categorization with Eye-tracking Paradigms
07:31

Defining the Role Of Language in Infants' Object Categorization with Eye-tracking Paradigms

Published on: February 8, 2019

6.5K

相关实验视频

Last Updated: May 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

465
Post-Movie Subliminal Measurement PMSM, for Investigating Implicit Social Bias
09:03

Post-Movie Subliminal Measurement PMSM, for Investigating Implicit Social Bias

Published on: February 29, 2020

5.7K
Defining the Role Of Language in Infants' Object Categorization with Eye-tracking Paradigms
07:31

Defining the Role Of Language in Infants' Object Categorization with Eye-tracking Paradigms

Published on: February 8, 2019

6.5K

科学领域:

  • 人工智能的人工智能
  • 自然语言处理自然语言处理.
  • 计算社会科学 计算社会科学

背景情况:

  • 大型语言模型 (LLM) 显示出令人印象深刻的功能,但通常会从训练数据中延续社会偏见.
  • 检测和量化LLM中的这些隐含偏见对于负责任的AI开发至关重要.

研究的目的:

  • 引入一种新的框架来检测LLMs中的社会偏见,将其作为一个假设测试问题.
  • 为了比较经典统计测试与贝叶斯推理对偏差量化的有效性.

主要方法:

  • 重构了偏见检测作为一个假设测试问题,零假设代表了隐含偏见的缺乏.
  • 利用二进制选择问题来测量各种LLM中的社会偏见 (例如,ChatGPT,DeepSeek-V3,Llama-3.1-70B).
  • 集成精确的二项式测试与贝叶斯推理,使用贝叶斯因子用于偏差检测和量化.

主要成果:

  • 与准确的二项式测试相比,贝叶斯因子在量化竞争假设的证据方面表现出更好的能力.
  • 贝叶斯因子对小样本大小具有稳定性,提供更可靠的偏差量化.
  • 在CrowS-Pairs数据集的英语和法语版本中,LLM偏见行为显示出一致性,其中一些细微的变化归因于社会文化背景.

结论:

  • 提出的假设测试框架,特别是贝叶斯因子,提供了一种强大的方法来检测和量化LLMs的社会偏见.
  • 贝叶斯推理在区分偏见的证据与无偏见的证据方面比经典测试提供了优势.
  • 偏见的跨语言一致性表明了潜在的模式,尽管文化细微差别需要进一步调查.