聊天机器人是可靠的文字注释器吗? 有时候有时候有时候有时候
Ross Deans Kristensen-McLachlan1,2, Miceal Canavan3, Marton Kárdos2
1Department of Linguistics, Cognitive Science, and Semiotics, Aarhus University, Aarhus 8000, Denmark.
PNAS nexus
|April 2, 2025
概括
与ChatGPT相比,开源大型语言模型 (LLM) 在社会科学文本注释方面表现不同. 监督模型,如DistilBERT,通常提供更可靠的结果,特别是在开放科学实践中.
科学领域:
- 社会科学 社会科学 社会科学
- 计算语言学 计算语言学
- 人工智能的人工智能
背景情况:
- 聊天GPT对社会科学文本注释有希望,但有缺点 (闭源,透明度,可重复性,成本,数据保护).
- 开源 (OS) 大型语言模型 (LLM) 为解决这些局限性提供了一个潜在的替代方案.
研究的目的:
- 系统地比较OS LLMs与ChatGPT和传统的监督机器学习分类器在文本注释任务中的性能.
- 评估零射击,少数射击学习和提示变化的对模型性能的影响.
主要方法:
- 对多个OS的LLM和ChatGPT进行比较评估.
- 利用零射击和少数射击学习与通用和自定义提示.
- 在美国新闻媒体推特的新数据集上测试模型,用于二进制文本注释.
- 与监督分类模型 (DistilBERT) 进行LLM性能比较.
主要成果:
- 在ChatGPT和各种OS LLM之间,在不同的注释任务中观察到显著的性能差异.
- 使用DistilBERT的监督分类器在总体上表现优于ChatGPT和评估的OS LLMs.
- 发现ChatGPT的性能对于实质性文本注释是不可靠的.
结论:
- 在社会科学研究中使用ChatGPT进行实质性文本注释时,建议谨慎使用,因为性能可变性和开放科学挑战.
- OS LLMs提供了一个替代方案,但需要仔细评估;像DistilBERT这样的监督方法仍然是一个强大的基准.
- 需要进一步的研究来优化OS LLMs的社会科学应用,并确保透明度和可重复性.
相关概念视频
Non-equilibrium in the Cell
4.1K
An important concept in studying metabolism and energy is that of chemical equilibrium. Most chemical reactions are reversible. They can proceed in both directions, releasing energy into their environment in one direction, and absorbing it from the environment in the other direction. The same is true for the chemical reactions involved in cell metabolism, such as the breaking down and building up of proteins into and from individual amino acids, respectively. Reactants within a closed system...
4.1K
Reliability and Validity
12.6K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.6K
Improving Translational Accuracy
2.5K
2.5K
Accuracy and Precision
8.6K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value. Highly accurate...
8.6K
Random and Systematic Errors
10.7K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
10.7K
Systematic Error: Methodological and Sampling Errors
1.4K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
1.4K


