Related Experiment Video
Updated: Sep 1, 2025

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
7.8K
Critical analysis on the reproducibility of visual quality assessment using deep features
Franz Götz-Hahn1, Vlad Hosu1, Dietmar Saupe1
1Dept. of Computer and Information Science, Universität Konstanz, Konstanz, Baden-Württemberg, Germany.
Plos One
|August 16, 2022
Summary
Supervised machine learning models require careful data splitting. This study reveals significant data leakage in image and video quality assessment, invalidating previously reported high performance results.
Area of Science:
- Computer Science
- Machine Learning
- Image Processing
- Video Processing
Background:
- Supervised machine learning models are typically trained on datasets partitioned into distinct training, validation, and testing subsets to ensure unbiased performance evaluation.
- The field of no-reference image and video quality assessment has recently seen reported performance metrics significantly exceeding established benchmarks.
- Concerns have been raised regarding the integrity of these high performance claims due to potential methodological flaws.
Purpose of the Study:
- To investigate and identify complex data leakage issues within the no-reference image and video quality assessment literature.
- To rigorously re-evaluate the performance of reported approaches after correcting for identified data leakage.
- To explore potential improvements through end-to-end variations of existing methods.
Main Methods:
- Analysis of common supervised machine learning data splitting practices (training, validation, test sets).
- Identification and characterization of various forms of test set information leakage into the training process.
- Re-evaluation of model performance on corrected datasets, comparing against state-of-the-art benchmarks.
- Investigation of end-to-end variations of the discussed quality assessment approaches.
Main Results:
- Complex data leakage, where information from the test set contaminates the training process, was identified in numerous studies.
- Claimed state-of-the-art performance results in the no-reference image and video quality assessment literature were found to be unattainable due to data leakage.
- Upon correction for data leakage, the performance of the analyzed approaches significantly decreased, falling substantially below existing benchmarks.
- End-to-end variations of the investigated methods did not yield improvements over the original approaches.
Conclusions:
- The prevalence of data leakage in no-reference image and video quality assessment research has led to inflated performance claims.
- Rigorous data handling and validation are crucial for reliable machine learning model evaluation, particularly in sensitive areas like perceptual quality assessment.
- Future research must prioritize robust methodologies to prevent data leakage and ensure the reproducibility and validity of results in this field.

