对三种语言模型的系统测试显示了语言准确性低,缺乏响应稳定性,以及对答案偏差
Vittoria Dentella1, Fritz Günther2, Evelina Leivada3,4
1Departament d'Estudis Anglesos i Alemanys, Universitat Rovira i Virgili, Tarragona 43002, Spain.
概括
像ChatGPT这样的大型语言模型 (LLM) 与人类相比显示不稳定和不准确的语法判断. 它们在语言任务上的表现揭示了偏见和缺乏可靠性,质疑它们作为人类语言模型的使用.
科学领域:
- 计算语言学计算语言学
- 认知科学是一种认知科学.
- 人工智能的人工智能是人工智能.
背景情况:
- 人类具有天生的语言判断能力.
- 大型语言模型 (LLM) 越来越多地被声称表现出类似人类的语言能力.
- 评估LLM区分语法和非语法句子的能力至关重要.
研究的目的:
- 在语言语法判断中评估著名的LLM (GPT-3变体,ChatGPT) 的稳定性和准确性.
- 将LLM绩效与人类判断模式进行比较.
- 调查八种不同的语言现象中的LLM反应.
主要方法:
- 每个模型的LLM在800个判断任务中进行了测试,涉及8种语言现象.
- 每个现象包括5个语法和5个非语法句子,重复10次.
- 人类参与者 (n=80) 完成了同样的任务进行比较.
主要成果:
- 对于语法句子,LLM表现出可变的机会之上准确性,而对于非语法句子,则表现出机会之下的准确性.
- 在语言现象中观察到LLM响应的显著不稳定性.
- 在所有测试的LLM中都发现了一致的"是-响应"偏差,重复并没有改善稳定性.
结论:
- 在识别 (非) 语法方面,LLM的表现与人类语言能力形成鲜明对比.
- 当前的LLM表现出显著的不稳定性和偏见,质疑它们作为人类语言学习模型的适用性.
- 需要进一步的研究来提高在语言处理任务中LLM的可靠性.
相关概念视频
Accuracy and Errors in Hypothesis Testing
201
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
201
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K
Systematic Error: Methodological and Sampling Errors
1.5K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
1.5K
Confirmation Biases
5.5K
The confirmation bias is the tendency to focus on information that confirms our existing beliefs and ignore information that is inconsistent with our expectations. For example, if you think that your professor is not very nice, you notice all of the instances of rude behavior exhibited by the professor while ignoring the countless pleasant interactions he is involved in on a daily basis. Have you ever fallen prey to the confirmation bias, either as the source or target of such bias?
5.5K
Errors In Hypothesis Tests
4.2K
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
4.2K


