Evaluating cognitive depth of AI-generated multiple-choice questions with Bloom's Taxonomy
Trang Thi Nguyen1, Linh Nguyen2, Ha Thi Nguyet Do2
1Faculty of Dentistry, Phenikaa University, Hanoi, Vietnam.
Plos One
|February 27, 2026
Summary
Large language models (LLMs) show strong performance in generating lower-level Bloom's Taxonomy multiple-choice questions (MCQs). Claude Sonnet 4 demonstrated superior alignment with higher-order cognitive skills in medical education.
Area of Science:
- Medical Education
- Artificial Intelligence in Education
- Cognitive Science
Background:
- Large language models (LLMs) are increasingly used for generating educational content, including multiple-choice questions (MCQs).
- The alignment of LLM-generated MCQs with established educational frameworks like Bloom's Taxonomy has not been thoroughly investigated.
- This study focuses on evaluating LLMs' ability to create MCQs across different cognitive levels within oral and maxillofacial anatomy.
Purpose of the Study:
- To assess the alignment of multiple widely used large language models (LLMs) with Bloom's Taxonomy cognitive levels.
- To compare the performance of different LLMs in generating MCQs across remembering, understanding, applying, analyzing, and evaluating/creating levels.
- To identify which LLMs best support higher-order thinking skills in medical education content generation.
Main Methods:
- Five leading LLMs (ChatGPT-4o, Copilot Pro, Claude Sonnet 4, Grok 3, DeepSeek R1) generated 300 MCQs from an oral and maxillofacial anatomy textbook.
- MCQs were designed to target five cognitive levels of Bloom's Taxonomy.
- Two independent investigators rated each MCQ using a 5-point Likert scale, with inter-rater reliability assessed via weighted Cohen's kappa.
Main Results:
- Inter-rater reliability was moderate to strong (kappa = 0.74–0.86).
- All LLMs performed well on lower cognitive levels (remembering, understanding, applying, evaluating/creating), with median scores above 4.
- Claude Sonnet 4 showed superior performance at higher cognitive levels (applying, analyzing, evaluating/creating) compared to other models.
- ChatGPT-4o, DeepSeek R1, and Grok 3 demonstrated a significant bias towards lower cognitive levels.
Conclusions:
- LLMs generally perform well in generating MCQs for lower cognitive levels of Bloom's Taxonomy.
- Claude Sonnet 4 exhibits the highest alignment with higher-order thinking skills among the evaluated LLMs.
- The findings suggest varying capabilities of LLMs in supporting comprehensive cognitive skill development in medical education.
More Related Videos
Related Concept Videos
Multiple Intelligences Theory
9.2K
Howard Gardner's theory of Multiple Intelligence proposes that there are nine distinct types of intelligence, each reflecting different ways of interacting with the world. Introduced in 1983 and expanded in subsequent years, Gardner's framework challenges the traditional notion of a single, generalized intelligence.
9.2K
Cognitive Learning
1.5K
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
1.5K
Critical Thinking II
5.1K
Critical thinking is a cognitive process with several attributes. The attributes of critical thinking include the following:
5.1K
Measures of Intelligence
8.7K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
8.7K
Introduction to Cognitive Psychology
2.8K
Cognitive psychology is the field of psychology dedicated to examining how people think. It attempts to explain how and why we think the way we do by studying the interactions among human thinking, emotion, creativity, language, and problem-solving, as well as other cognitive processes. Cognitive psychology studies how information is processed and manipulated in remembering, thinking, and knowing.
This field emerged in the mid-20th century, following a period dominated by behaviorism, which...
This field emerged in the mid-20th century, following a period dominated by behaviorism, which...
2.8K
Lazarus's Cognitive Appraisal Theory
2.3K
Cognitive psychologist Richard Lazarus proposed the cognitive-mediational theory of emotions, which emphasizes how individuals' assessments of stressors significantly affect their experience of stress. According to Lazarus, the stress response is determined by a two-step appraisal process: primary appraisal and secondary appraisal. These cognitive appraisals help individuals evaluate the potential impact of a stressor and determine the adequacy of their coping resources.
Primary Appraisal:...
Primary Appraisal:...
2.3K


