Related Experiment Video
Updated: Mar 9, 2026

The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups
Published on: May 13, 2022
Evaluating LLMs' divergent thinking capabilities for scientific idea generation with minimal context.
Kai Ruan1, Xuan Wang2, Jixiang Hong1
1Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China.
Large Language Models (LLMs) show promise in science, but current tests miss their idea generation skills. A new benchmark, LiveIdeaBench, reveals LLM creativity is not well predicted by general intelligence scores.
Area of Science:
- Artificial Intelligence
- Computational Science
- Scientific Discovery
Background:
- Large Language Models (LLMs) excel at scientific tasks like literature analysis and experimental design.
- Existing benchmarks often use rich contextual inputs, potentially overlooking specific creative capabilities.
- Evaluating divergent thinking in LLMs for scientific idea generation remains a challenge.
Purpose of the Study:
- Introduce LiveIdeaBench, a novel benchmark for assessing LLMs' scientific idea generation.
- Evaluate LLMs' divergent thinking using single-keyword prompts, inspired by creativity theory.
- Assess the relationship between general intelligence scores and scientific idea generation capabilities.
Main Methods:
- Developed LiveIdeaBench, a benchmark using 1180 keywords across 22 scientific domains.
- Employed a dynamic panel of over 40 state-of-the-art LLMs.
- Assessed generated ideas across five dimensions: originality, feasibility, fluency, flexibility, and clarity.
Main Results:
- LLMs' scientific idea generation capabilities are poorly predicted by standard general intelligence metrics.
- Models with lower general intelligence scores can exhibit comparable creative performance to higher-scoring models.
- Significant variation exists in creative idea generation across different LLMs.
Conclusions:
- Current evaluation methods may not fully capture LLMs' potential for scientific creativity.
- Specialized benchmarks like LiveIdeaBench are crucial for understanding and improving LLM scientific idea generation.
- Developing LLMs for scientific idea generation may necessitate distinct training strategies compared to general problem-solving.
More Related Videos
10:26Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
06:45Task Interruption and Resumption Paradigm for Testing the Activation and Pursuit of an Abstract Thinking Goal
Published on: April 18, 2017
Related Concept Videos
Creative Thinking
Divergent thinking is the...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Counterfactual Thinking
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Ampere-Maxwell's Law: Problem-Solving
To solve the problem, we can use the equations from the analysis of an RC circuit and Maxwell's version of Ampère's law.
For the first part of the...
High-Level and Low-Level Awareness