作为研究科学家的ChatGPT:探讨GPT作为研究图书馆员,研究伦理学家,数据生成器和数据预测器的能力
Steven A Lehr1, Aylin Caliskan2, Suneragiri Liyanage3
1Cangrade, Inc., Watertown, MA 02472.
概括
聊天GPT显示出作为研究工具的潜力,作为伦理学家和数据模拟器出色,但在准确引用和预测新型科学数据方面扎.
科学领域:
- 心理学科学研究 心理学科学研究
- 人工智能能力的人工智能能力
- 研究中的大型语言模型.
背景情况:
- 人工智能 (AI) 在科学研究中的整合正在迅速发展.
- 评估像ChatGPT (GPT-3.5和GPT-4) 这样的先进AI模型的能力,对于理解它们在科学应用中的潜力和局限性至关重要.
研究的目的:
- 系统地评估GPT-3.5和GPT-4在关键科学研究组件中的表现.
- 评估AI作为心理学科学中的研究图书馆员,伦理学家,数据生成器和新型数据预测器的实用性.
主要方法:
- 为了探讨人工智能的能力,进行了四项研究:研究图书馆员 (参考生成),研究伦理学家 (伦理协议分析),数据生成器 (模拟已知的模式) 和新数据预测器 (预测新的经验结果).
- 心理学科学作为测试的领域,使用虚构的研究协议和分析AI产生的输出与既定的基准.
主要成果:
- 在GPT模型中,参考的幻觉率显著 (GPT-3.5: 36.0%,GPT-4: 5.4%),GPT-4在识别错误方面有所改善.
- 在研究协议中,GPT-4在识别伦理违规 (p-hacking) 方面表现出熟练 (88.6%是公开的,72.6%是微妙的),而GPT-3.5则没有.
- 两种模型在生成数据时都复制了已知的文化偏见,表明模拟现有发现的能力.
- 无论是GPT-3.5还是GPT-4,都没有成功地预测出超出其培训数据的新型研究结果.
结论:
- 聊天GPT (GPT-3.5和GPT-4) 是一个有缺陷但正在改进的研究图书馆员,容易产生虚构的引用.
- 作为研究伦理学家,GPT-4显示出有前途的潜力,能够在科学协议中检测出伦理问题.
- 这些模型可以生成模拟已知的模式的数据,这对于模拟结果很有用,但对于发现新的结果并不有用.
- 目前的人工智能模型在预测新的经验数据方面存在局限性,突出了科学AI应用未来发展的领域.
相关概念视频
Archival Research
16.0K
Some researchers gain access to large amounts of data without interacting with a single research participant. Instead, they use existing records to answer various research questions. This type of research approach is known as archival research. Archival research relies on looking at past records or data sets to look for interesting patterns or relationships. For example, a researcher might access the academic records of all individuals who enrolled in college within the past ten years and...
16.0K
Ethics in Research
22.9K
Today, scientists agree that good research is ethical in nature and is guided by a basic respect for human dignity and safety. However, this has not always been the case. Modern researchers must demonstrate that the research they perform is ethically sound.
22.9K


