Related Experiment Video
Updated: Jun 4, 2026

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
General scales unlock AI evaluation with explanatory and predictive power
Lexin Zhou1,2,3,4, Lorenzo Pacchiardi5, Fernando Martínez-Plumed6
1Princeton University, Princeton, NJ, USA. lz5066@princeton.edu.
Nature
|April 1, 2026
Summary
New AI evaluation scales predict performance across tasks by profiling AI abilities and demands. This approach enhances understanding and reliable deployment of artificial intelligence (AI) systems.
Area of Science:
- Artificial Intelligence
- Machine Learning Evaluation
- AI Benchmarking
Background:
- Current artificial intelligence (AI) benchmarking lacks explanatory and predictive power for general-purpose systems.
- Limited transferability across tasks hinders understanding of AI capabilities.
- Existing methods struggle to predict AI performance on novel tasks.
Purpose of the Study:
- Introduce general scales for AI evaluation to elicit demand and ability profiles.
- Quantify general strengths and limits of AI systems.
- Robustly predict AI performance on new task instances.
Main Methods:
- Developed a fully automated methodology using 18 rubrics.
- Captured a broad range of cognitive and intellectual demands.
- Applied scales to 15 large language models (LLMs) and 63 tasks.
Main Results:
- Demand and ability profiles provide insights into benchmark construct validity.
- Explained conflicting claims regarding AI reasoning capabilities.
- Achieved high instance-level predictive power, outperforming black-box predictors, especially in out-of-distribution settings.
Conclusions:
- The general scales offer a foundation for a science of AI evaluation.
- Enables superior AI performance prediction for new tasks and benchmarks.
- Underpins the reliable deployment of AI systems.