Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 4, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

General scales unlock AI evaluation with explanatory and predictive power.

Lexin Zhou1,2,3,4, Lorenzo Pacchiardi5, Fernando Martínez-Plumed6

  • 1Princeton University, Princeton, NJ, USA. lz5066@princeton.edu.

Nature
|April 1, 2026
PubMed

Related Concept Videos

Variation01:19

Variation

8.3K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
8.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Efficacy and Safety of Baxdrostat in the Management of Uncontrolled Hypertension: A Systematic Review and Meta-Analysis.

American journal of hypertension·2026
Same author

Revisiting Rogers' Paradox in the context of human-AI interaction.

Philosophical transactions. Series A, Mathematical, physical, and engineering sciences·2026
Same author

Learning patient-specific spatial biomarker dynamics via operator learning for Alzheimer's disease progression.

NPJ systems biology and applications·2026
Same author

The triangular drivers of bone aging: mechanistic insights and therapeutic targets in cellular senescence, estrogen deficiency, and gut microenvironment dysregulation.

Frontiers in cell and developmental biology·2026
Same author

Corrigendum to "Integrated amplification of NADPH-regenerating modules enhances cytidine biosynthesis in <i>Escherichia coli</i>" [Synth Syst Biotechnol 12 (2026) 320-329].

Synthetic and systems biotechnology·2026
Same author

Interfacial Characteristics of HgCdTe Infrared Detectors Grown on Alternative Substrates.

Sensors (Basel, Switzerland)·2026
Summary

New AI evaluation scales predict performance across tasks by profiling AI abilities and demands. This approach enhances understanding and reliable deployment of artificial intelligence (AI) systems.

Area of Science:

  • Artificial Intelligence
  • Machine Learning Evaluation
  • AI Benchmarking

Background:

  • Current artificial intelligence (AI) benchmarking lacks explanatory and predictive power for general-purpose systems.
  • Limited transferability across tasks hinders understanding of AI capabilities.
  • Existing methods struggle to predict AI performance on novel tasks.

Purpose of the Study:

  • Introduce general scales for AI evaluation to elicit demand and ability profiles.
  • Quantify general strengths and limits of AI systems.
  • Robustly predict AI performance on new task instances.

Main Methods:

  • Developed a fully automated methodology using 18 rubrics.
  • Captured a broad range of cognitive and intellectual demands.

Related Experiment Videos

Last Updated: Jun 4, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

  • Applied scales to 15 large language models (LLMs) and 63 tasks.
  • Main Results:

    • Demand and ability profiles provide insights into benchmark construct validity.
    • Explained conflicting claims regarding AI reasoning capabilities.
    • Achieved high instance-level predictive power, outperforming black-box predictors, especially in out-of-distribution settings.

    Conclusions:

    • The general scales offer a foundation for a science of AI evaluation.
    • Enables superior AI performance prediction for new tasks and benchmarks.
    • Underpins the reliable deployment of AI systems.