Related Experiment Video
Updated: Nov 28, 2025

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Understanding Deep Learning Performance through an Examination of Test Set Difficulty: A Psychometric Case Study
John P Lalor1, Hao Wu2, Tsendsuren Munkhdalai3
1College of Information and Computer Sciences, University of Massachusetts, Amherst.
Abstract:
Interpreting the performance of deep learning models beyond test set accuracy is challenging. Characteristics of individual data points are often not considered during evaluation, and each data point is treated equally. We examine the impact of a test set question's difficulty to determine if there is a relationship between difficulty and performance. We model difficulty using well-studied psychometric methods on human response patterns. Experiments on Natural Language Inference (NLI) and Sentiment Analysis (SA) show that the likelihood of answering a question correctly is impacted by the question's difficulty. As DNNs are trained with more data, easy examples are learned more quickly than hard examples.
More Related Videos
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
06:11High-definition Transcranial Direct Current Stimulation over Right Dorsolateral Prefrontal Cortex to Enhance Metacognitive Sensitivity
Published on: September 26, 2025
Related Concept Videos
Theory of Attribution II: Kelley's Covariation Theory
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...