Related Experiment Video
Updated: Jun 4, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
463
Benchmarking the speed-accuracy tradeoff in object recognition by humans and neural networks
Ajay Subramanian1,2, Sara Price3,4, Omkar Kumbhar5,6
1Department of Psychology, New York University, New York, NY, USA.
Journal of Vision
|January 3, 2025
Summary
Human object recognition involves a speed-accuracy tradeoff (SAT). This study introduces a new dataset to compare human SAT with dynamic neural networks, finding cascaded networks best model this crucial cognitive skill.
Area of Science:
- Cognitive Science
- Computer Vision
- Artificial Intelligence
Background:
- Active object recognition is vital for real-world tasks like driving and reading.
- Human performance demonstrates a flexible speed-accuracy tradeoff (SAT), a critical but computationally challenging skill.
- Existing computational models often fail to adequately incorporate the temporal dynamics of decision-making.
Purpose of the Study:
- To introduce the first dataset for studying the speed-accuracy tradeoff (SAT) in ImageNet object recognition.
- To compare human SAT performance with various dynamic neural network architectures.
- To identify neural network models that best capture human temporal decision-making in object recognition.
Main Methods:
- Collected data from 148 human observers performing a time-sensitive 16-way ImageNet categorization task.
- Developed dynamic neural networks that adapt computation based on available inference time, using repetition count (layers, recurrent cycles, early exits) as a time analog.
- Compared human and network performance using metrics like SAT-fit error, category-wise correlation, and SAT-curve steepness.
Main Results:
- Human accuracy in object recognition increases with reaction time, confirming the expected speed-accuracy tradeoff.
- Cascaded dynamic neural networks showed the strongest correlation with human SAT performance across multiple metrics.
- Convolutional recurrent networks, previously favored for modeling human object recognition, performed poorly in this SAT benchmark.
Conclusions:
- Time is a critical, scarce resource in human object recognition, necessitating models that account for temporal adaptation.
- Cascaded dynamic neural networks offer a promising computational framework for modeling the human speed-accuracy tradeoff in visual recognition.
- The study highlights the limitations of traditional recurrent network models in capturing dynamic temporal decision-making processes.

