AIのエコー: LLMの出力におけるプロット多様性の欠如を定量化
Weijia Xu1, Nebojsa Jojic1, Sudha Rao1
1Microsoft Research, Redmond, WA 98052.
まとめ
GPT-4やLLaMA-3のような大規模な言語モデル (LLM) は,繰り返し物語のプロットを生成し,集団的な創造性を制限します. AIによって生成されたコンテンツのオリジナリティの欠如を 定量化しています
科学分野:
- 人工知能
- 自然言語処理
- コンピューターの創造性
背景:
- 大型言語モデル (LLM) は,クリエイティブなコンテンツの生成にますます使用されています.
- 重要な質問は,LLMが多種多様なアイデアを通して真の集団的な創造性を育むことができるかどうかです.
研究 の 目的:
- 最先端のLLM (GPT-4,LLaMA-3) で生成されたプロット要素の多様性を評価する.
- 土地の独自性を測定するための自動メトリックであるSui Generisスコアを導入し,検証する.
主な方法:
- 同様のプロンプトを使用して,GPT-4とLLaMA-3からストーリー生成を検証した.
- 複数のLLM世代にわたるプロット要素のユニークさを定量化するためにSui Generisスコアを開発しました.
- Sui Generisのスコアと感知された驚きのレベルを相関させるためのヒトの評価を実施しました.
主要な成果:
- LLMによって生み出された物語は,世代やモデルにまたがるプロット要素を頻繁に備えています.
- 人間が書いた物語は LLM の出力と比較して 劇情要素がかなり独特です
- スイ・ジェネリスのスコアは 驚異に対する人間の判断と 適度な相関を示し 効果を証明しています
結論:
- 現在のLLMは,繰り返しのプロット要素の生成のために,集合的な創造性を十分に強化しないかもしれません.
- Sui Generisのスコアは AIによって生成された物語のオリジナリティを 評価するための 信頼性の高い自動的な方法を提供します
- クリエイティブな応用のためのLLMの多様性を高めるためにさらなる研究が必要です.
さらに関連する動画
関連する概念動画
Quantifying and Rejecting Outliers: The Grubbs Test
2.0K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.0K
Variability: Analysis
189
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
189
Improving Translational Accuracy
2.7K
2.7K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Distributions to Estimate Population Parameter
4.3K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.3K
Language Development
444
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
444


