Related Experiment Video
Updated: Jun 17, 2026

Using the Visual World Paradigm to Study Sentence Comprehension in Mandarin-Speaking Children with Autism
Published on: October 3, 2018
The performance of ChatGPT and Bing on a computerized adaptive test of verbal intelligence
Balázs Klein1, Kristof Kovacs2
1Testar Ltd., Budapest, Hungary.
Abstract:
We administered a computerized adaptive test of vocabulary three times to assess the verbal intelligence of chatGPT (GPT 3.5) and Bing (based on GPT 4). There was no difference between their performance; both performed at a high level, outperforming approximately 95% of humans and scoring above the level of native speakers with a doctoral degree. In 42% of test items that were administered more than once these large language models provided different answers to the same question in different sessions. They never engaged in guessing, but provided hallucinations: answers that were not among the options. Such hallucinations were not triggered by the inability to answer correctly as the same questions evoked correct answers in other sessions. The results implicate that psychometric tools developed for humans have limitations when assessing AI, but they also imply that computerised adaptive testing of verbal ability is an appropriate tool to critically evaluate the performance of large language models.
Related Concept Videos
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Binet's Contribution to Measures of Intelligence

