Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Multiple Comparison Tests01:13

Multiple Comparison Tests

4.0K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.0K
Humanistic Psychology01:24

Humanistic Psychology

1.6K
Humanistic psychology emerged in the mid-20th century as a response to the deterministic and pessimistic nature of behaviorism and psychoanalysis. While behaviorism focused on observable behaviors influenced by the environment and psychoanalysis delved into unconscious motivations, both theories suggested that human actions lacked free will. In contrast, humanistic psychology offers a perspective that emphasizes the innate potential for goodness and growth within every individual.
This approach...
1.6K
Mechanical Efficiency of Real Machines01:14

Mechanical Efficiency of Real Machines

894
The mechanical efficiency of a machine is a fundamental concept that describes how effectively a machine can convert input work into output work. According to this concept, the efficiency of a machine is equal to the ratio of the output work to the input work. An ideal machine, meaning a machine that has no energy losses, has an efficiency of one. This implies that the input work and the output work are equal.
However, in reality, no machine can be truly ideal, and all of them experience some...
894

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Connectome quality converges predictably to reveal optimal stopping points during proofreading.

bioRxiv : the preprint server for biology·2026
Same author

Detecting dataset bias in medical AI using a generalized and modality agnostic auditing approach.

NPJ digital medicine·2026
Same author

Physical contact reveals a hidden layer of cortical architecture.

bioRxiv : the preprint server for biology·2026
Same author

Learned adaptive properties for mitigation of weight perturbations in embedded spiking networks.

Frontiers in neuroscience·2026
Same author

3D Neuromodulation in Neural Organoids with Shell MEAs.

Advanced healthcare materials·2026
Same author

EM and XRM Connectomics Imaging and Experimental Metadata Standards.

ArXiv·2026

Related Experiment Video

Updated: Sep 28, 2025

One Dimensional Turing-Like Handshake Test for Motor Intelligence
14:05

One Dimensional Turing-Like Handshake Test for Motor Intelligence

Published on: December 15, 2010

27.8K

A framework for rigorous evaluation of human performance in human and machine learning comparison studies.

Hannah P Cowley1, Mandy Natter2, Karla Gray-Roncal2

  • 1The Johns Hopkins University Applied Physics Laboratory, Research and Exploratory Development Department, Laurel, MD, 20723, USA. Hannah.Cowley@jhuapl.edu.

Scientific Reports
|April 1, 2022
PubMed
Summary

A new framework standardizes human performance evaluation for machine learning comparisons. This ensures reliable assessments of artificial intelligence capabilities against human cognition, advancing AI research.

More Related Videos

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

703
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

2.6K

Related Experiment Videos

Last Updated: Sep 28, 2025

One Dimensional Turing-Like Handshake Test for Motor Intelligence
14:05

One Dimensional Turing-Like Handshake Test for Motor Intelligence

Published on: December 15, 2010

27.8K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

703
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

2.6K

Area of Science:

  • Artificial Intelligence
  • Cognitive Science
  • Human-Computer Interaction

Background:

  • Direct comparisons between human and machine learning (ML) performance are crucial for validating AI.
  • The ML community lacks a standardized framework for evaluating human performance in comparative studies.
  • Existing evaluation methods often overlook fundamental differences between human and algorithmic cognition.

Purpose of the Study:

  • To address the lack of a standardized framework for human performance evaluation in ML comparisons.
  • To propose guiding principles for designing robust human evaluation studies.
  • To enhance the accuracy and reproducibility of human-AI performance comparisons.

Main Methods:

  • Demonstration of common pitfalls in human performance evaluation design.
  • Proposal of a standardized framework with three key principles.
  • Illustration of the framework's application using a one-shot learning task study.

Main Results:

  • Identified critical considerations for designing human evaluations, including understanding cognitive differences.
  • Emphasized trial matching between human participants and algorithms.
  • Advocated for adopting best practices from psychology research, including supplementary data collection and ethical protocols.

Conclusions:

  • The proposed framework offers a standardized approach for evaluating human performance against ML algorithms.
  • Adoption of this framework can improve the reliability and reproducibility of AI performance comparisons.
  • This standardization is vital for advancing the field of artificial intelligence.