评价大型语言模型的性能和可靠性,引用和参考文献在学术写作:跨学科研究研究
Joseph Mugaanyi1, Liuying Cai2, Sumei Cheng2
1Department of Hepato-Pancreato-Biliary Surgery, Ningbo Medical Center Lihuili Hospital, Health Science Center, Ningbo University, Ningbo, China.
Journal of medical Internet research
|April 5, 2024
概括
大型语言模型 (LLM) 在各个学术领域显示出不同的引用准确性. 聊天GPT (GPT-3.5) 在自然科学领域产生了比人文领域更准确的数字物体识别器 (DOI).
科学领域:
- 科学研究科学研究
- 学术写作学术写作
- 数字对象识别器 (DOI) 是一个数字对象识别器.
背景情况:
- 大型语言模型 (LLM) 在学术界迅速获得突出地位.
- 像ChatGPT (GPT-3.5) 这样的LLM在学术任务中的能力需要彻底评估.
研究的目的:
- 评估ChatGPT (GPT-3.5) 产生的引文和引用的准确性.
- 为了比较ChatGPT在自然科学和人文学科之间的引文生成方面的表现.
主要方法:
- 研究人员促使ChatGPT生成带有引用的介绍部分.
- 两个研究人员独立验证了生成的引用和数字对象识别器 (DOI) 的准确性.
- 自然科学和人文科学之间的引用和DOI准确性进行了比较.
主要成果:
- 聊天GPT在10个主题 (5个自然科学,5个人文科学) 中产生了102个引用.
- 引用存在率相似 (72.7%的自然科学与76.6%的人文科学).
- 数字物体识别器 (DOI) 的存在 (70.9%对38.3%) 和准确性 (32.7%对8.5%) 在自然科学中显著更高,在人文学科中更高的DOI幻觉.
结论:
- 聊天GPT的引用和引用生成准确性在学科之间有很大的差异.
- 数字物体识别器 (DOI) 标准和细微差别中的学科变化会影响LLM的表现.
- 研究人员必须批判性地评估AI写作工具的引用准确性,并考虑特定领域的AI模型.
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Improving Translational Accuracy
2.6K
2.6K
Self-Evaluation: Self-Enhancement and Self-Verification
5.2K
Social psychologists have documented that feeling good about ourselves and maintaining positive self-esteem is a powerful motivator of human behavior (Tavris & Aronson, 2008). In the United States, members of the predominant culture typically think very highly of themselves and view themselves as good people who are above average on many desirable traits (Ehrlinger, Gilovich, & Ross, 2005). Often, our behavior, attitudes, and beliefs are affected when we experience a threat to our...
5.2K
Accuracy and Precision
8.8K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value. Highly accurate...
8.8K
Sample Size Calculation
3.3K
Knowledge of the sample size is the first requirement to conduct random sampling or an experiment. The sample size is the total number of units, observations, or groups (in some cases) used to get the data to estimate a population parameter. As the name suggests, the sample size is that of the sample drawn from the population and differs from the population size.
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
3.3K
Systematic Error: Methodological and Sampling Errors
1.5K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
1.5K


