聊天GPT和Bing在言语智能的计算机自适应测试中的表现
Balázs Klein1, Kristof Kovacs2
1Testar Ltd., Budapest, Hungary.
PloS one
|July 25, 2024
概括
像ChatGPT和Bing这样的大型语言模型显示出高口语智能,表现优于大多数人类. 然而,他们不一致地回答问题,并表现出幻觉,表明当前AI评估工具的局限性.
科学领域:
- 人工智能的人工智能
- 心理测量 心理测量 心理测量
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 展示了先进的能力,需要强大的评估方法.
- 使用以人为中心的心理测量工具评估人工智能系统的口头智能存在独特的挑战.
- 以前的研究已经探讨了AI的性能,但往往缺乏标准化的口头智能评估.
研究的目的:
- 用计算机自适应测试来评估ChatGPT (GPT-3.5) 和Bing (GPT-4) 的语言智力.
- 识别LLM表现中的不一致性和奇特的响应模式,如幻觉.
- 评估人类心理测量工具的适用性,以评估先进的AI语言模型.
主要方法:
- 对ChatGPT和Bing进行了三次计算机化适应性词汇测试.
- 将AI性能与人类基准进行比较,特别是拥有博士学位的母语人士.
- 在重复的测试中分析了反应的一致性和非答案选项 (幻觉) 的发生.
主要成果:
- 聊天GPT和Bing都表现出很高的水平,超过了大约95%的人类表现.
- 在ChatGPT和Bing的整体口头智力得分之间没有发现显著差异.
- 在42%的重复项目中,LLMs为同一个问题提供了不同的答案,并且在没有猜测的情况下表现出幻觉 (提供不存在的选项).
结论:
- 人类心理测量工具在评估AI时存在局限性,特别是在响应一致性和幻觉检测方面.
- 计算机自适应测试是一种可行的方法,用于批判性地评估大型语言模型的语言能力.
- 法律学的表现凸显了开发AI特定评估方法的必要性,以确保可靠和有效的评估.
相关概念视频
Measures of Intelligence
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Binet's Contribution to Measures of Intelligence
Alfred Binet, along with his student Théophile Simon, was tasked by the French Ministry of Education in 1904 to create a method for identifying students who struggled to learn through conventional classroom instruction. This initiative aimed to address overcrowding by placing such students in specialized schools. Binet and Simon developed an intelligence test comprising 30 tasks, ranging from simple commands, like touching one's nose or ear, to more complex tasks, such as drawing designs from...


