Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Measures of Intelligence01:29

Measures of Intelligence

Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Wechsler's Contribution to Measures of Intelligence01:23

Wechsler's Contribution to Measures of Intelligence

David Wechsler, a psychologist who worked with World War I veterans, developed a significant IQ test in 1939 called the Wechsler-Bellevue Intelligence Scale. This test was innovative because it combined several subtests that measured both verbal and nonverbal skills, reflecting Wechsler's belief that intelligence is a global capacity involving purposeful action, rational thinking, and effective interaction with the environment. This test later evolved into the Wechsler Adult Intelligence Scale...
Binet's Contribution to Measures of Intelligence01:23

Binet's Contribution to Measures of Intelligence

Alfred Binet, along with his student Théophile Simon, was tasked by the French Ministry of Education in 1904 to create a method for identifying students who struggled to learn through conventional classroom instruction. This initiative aimed to address overcrowding by placing such students in specialized schools. Binet and Simon developed an intelligence test comprising 30 tasks, ranging from simple commands, like touching one's nose or ear, to more complex tasks, such as drawing designs from...
Triarchic Theory of Intelligence01:24

Triarchic Theory of Intelligence

Robert Sternberg's triarchic theory of intelligence posits that intelligence is composed of three distinct but interrelated components: analytical, creative, and practical intelligence.
Higher Mental Functions of the Brain: Language01:10

Higher Mental Functions of the Brain: Language

Language is a system of communication that allows the expression of thoughts, ideas, and feelings. The brain processes language in both hemispheres.
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Cattell's Theory of Intelligence01:25

Cattell's Theory of Intelligence

Raymond Cattell, along with John Horn, made significant contributions to our understanding of intelligence by distinguishing between two types: fluid intelligence and crystallized intelligence.
Fluid intelligence involves the capacity to solve new problems and adapt to unfamiliar situations. It's the type of intelligence individuals use when they encounter a novel problem or puzzle that requires innovative thinking. For instance, figuring out how to operate a new gadget relies heavily on fluid...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The Role of Inhibitory Control in Moderating Working Memory Training Transfer: Evidence from a Preregistered Trial with Older Adults.

Research square·2026
Same author

The Critical Period Microbiota Shape Brain Plasticity.

bioRxiv : the preprint server for biology·2026
Same author

Thermal Ablation with or without Liver Transplantation for Hepatocellular Carcinoma in Older Adults with Child-Pugh Class A Cirrhosis.

Journal of vascular and interventional radiology : JVIR·2026
Same author

Unsupervised clustering FLIM-phasor from multifunctionalized nanoparticles in living cancer cells.

Methods and applications in fluorescence·2026
Same author

The nBAF complex subunit CREST/SS18L1 regulates hippocampal memory processes via tyrosine 397 and histone acetyltransferase CBP.

Cell reports·2026
Same author

Circadian reprogramming by timed sodium intake reveals transcriptional pathways of daily salt handling in the colon.

Science advances·2026

Related Experiment Video

Updated: May 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Evaluating the Intelligence of large language models: A comparative study using verbal and visual IQ tests.

Sherif Abdelkarim1, David Lu1, Dora-Luz Flores2

  • 1University of California Irvine, 510 E Peltason Dr., Irvine, 92617, CA, USA.

Computers in Human Behavior. Artificial Humans
|May 22, 2026
PubMed
Summary

Large language models (LLMs) show varied reasoning skills, excelling in verbal tasks but struggling with numerical and abstract arithmetic. IQ tests offer a benchmark for evaluating LLM intelligence and tracking progress over time.

Keywords:
Artificial IntelligenceIntelligence QuotientLarge language models

More Related Videos

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
08:05

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques

Published on: June 30, 2020

Related Experiment Videos

Last Updated: May 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
08:05

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques

Published on: June 30, 2020

Area of Science:

  • Artificial Intelligence
  • Cognitive Science
  • Machine Learning Evaluation

Background:

  • Large language models (LLMs) demonstrate proficiency on specialized tasks, but their general reasoning capabilities require further investigation.
  • Existing benchmarks may not fully capture the breadth of cognitive abilities in advanced AI systems.

Purpose of the Study:

  • To evaluate the general reasoning abilities of 18 diverse large language models using a comprehensive IQ test suite.
  • To analyze the impact of model scale, multimodality, and multi-agent reflection on reasoning performance.
  • To establish IQ tests as a standardized benchmark for longitudinal comparison of LLM cognitive abilities against human norms.

Main Methods:

  • Administered a 14-section IQ suite covering verbal, numerical, and visual reasoning tasks to 18 LLMs.
  • Implemented a multi-agent reflection variant where models critique and revise answers.
  • Analyzed performance variations based on model size, multimodality, and reasoning task type.

Main Results:

  • Observed a significant bias towards verbal reasoning (e.g., GPT-4: 79% accuracy) over numerical reasoning (53% accuracy).
  • Identified a pronounced modality gap, with text-based IQ scores (≈125) significantly higher than visual-based scores (≈103).
  • Found persistent difficulties in abstract arithmetic tasks (≤20% accuracy) and modest gains from multi-agent reflection in frontier models.

Conclusions:

  • LLM intelligence, as measured by IQ tests, scales with model size but exhibits non-uniform gains across different reasoning domains.
  • IQ tests provide a valuable, human-referenced framework for evaluating and comparing LLM cognitive abilities, despite limitations.
  • Further research is needed to develop more comprehensive AI evaluation benchmarks that capture nuanced reasoning capabilities.