人工智能中的回声:量化LLM输出中的缺陷多样性
Weijia Xu1, Nebojsa Jojic1, Sudha Rao1
1Microsoft Research, Redmond, WA 98052.
概括
像GPT-4和LLaMA-3这样的大型语言模型产生重复的故事情节,限制了集体创造力. 一个新的指标,Sui Generis得分, 量化了人工智能产生的内容缺乏原创性.
科学领域:
- 人工智能
- 自然语言处理
- 计算创造力
背景情况:
- 大型语言模型 (LLM) 越来越多地用于创意内容生成.
- 一个关键的问题是,LLM能否通过多样化的创意来促进真正的集体创造力.
研究的目的:
- 通过最先进的LLM (GPT-4,LLaMA-3) 在故事生成中评估情节元素的多样性.
- 引入和验证一个自动指标,即Sui Generis评分,用于测量图片的独特性.
主要方法:
- 从GPT-4和LLaMA-3使用相同的提示生成故事.
- 开发了Sui Generis评分来量化多个LLM代的情节元素的独特性.
- 进行人体评估以将Sui Generis的分数与感知到的惊喜水平相关联.
主要成果:
- 通过LLM生成的故事经常具有跨世代和模式的故事情节元素.
- 与LLM产品相比,人类撰写的故事具有更多独特的情节元素.
- 苏格兰的Sui Generis得分与人类的惊喜判断有中等的相关性, 证实了它的有效性.
结论:
- 由于重复的情节元素生成, 目前的LLM可能不足以促进集体创造力.
- 提供可靠的自动方法来评估人工智能所产生的故事的原创性.
- 需要进一步的研究来增强创意应用的LLM多样性.
相关概念视频
Quantifying and Rejecting Outliers: The Grubbs Test
2.0K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.0K
Variability: Analysis
189
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
189
Improving Translational Accuracy
2.7K
2.7K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Distributions to Estimate Population Parameter
4.3K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.3K
Language Development
444
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
444


