Related Experiment Video
Updated: May 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Intelligence of large language models: A comparative study using verbal and visual IQ tests
Sherif Abdelkarim1, David Lu1, Dora-Luz Flores2
1University of California Irvine, 510 E Peltason Dr., Irvine, 92617, CA, USA.
Large language models (LLMs) show varied reasoning skills, excelling in verbal tasks but struggling with numerical and abstract arithmetic. IQ tests offer a benchmark for evaluating LLM intelligence and tracking progress over time.
Area of Science:
- Artificial Intelligence
- Cognitive Science
- Machine Learning Evaluation
Background:
- Large language models (LLMs) demonstrate proficiency on specialized tasks, but their general reasoning capabilities require further investigation.
- Existing benchmarks may not fully capture the breadth of cognitive abilities in advanced AI systems.
Purpose of the Study:
- To evaluate the general reasoning abilities of 18 diverse large language models using a comprehensive IQ test suite.
- To analyze the impact of model scale, multimodality, and multi-agent reflection on reasoning performance.
- To establish IQ tests as a standardized benchmark for longitudinal comparison of LLM cognitive abilities against human norms.
Main Methods:
- Administered a 14-section IQ suite covering verbal, numerical, and visual reasoning tasks to 18 LLMs.
- Implemented a multi-agent reflection variant where models critique and revise answers.
- Analyzed performance variations based on model size, multimodality, and reasoning task type.
Main Results:
- Observed a significant bias towards verbal reasoning (e.g., GPT-4: 79% accuracy) over numerical reasoning (53% accuracy).
- Identified a pronounced modality gap, with text-based IQ scores (≈125) significantly higher than visual-based scores (≈103).
- Found persistent difficulties in abstract arithmetic tasks (≤20% accuracy) and modest gains from multi-agent reflection in frontier models.
Conclusions:
- LLM intelligence, as measured by IQ tests, scales with model size but exhibits non-uniform gains across different reasoning domains.
- IQ tests provide a valuable, human-referenced framework for evaluating and comparing LLM cognitive abilities, despite limitations.
- Further research is needed to develop more comprehensive AI evaluation benchmarks that capture nuanced reasoning capabilities.
Related Concept Videos
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Wechsler's Contribution to Measures of Intelligence
Binet's Contribution to Measures of Intelligence
Triarchic Theory of Intelligence
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Cattell's Theory of Intelligence
Fluid intelligence involves the capacity to solve new problems and adapt to unfamiliar situations. It's the type of intelligence individuals use when they encounter a novel problem or puzzle that requires innovative thinking. For instance, figuring out how to operate a new gadget relies heavily on fluid...
