Related Experiment Video
Updated: Jul 4, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
572
Automated Quality Evaluation of Large-Scale Benchmark Datasets for Vision-Language Tasks
Ruibin Zhao1,2, Zhiwei Xie1, Yipeng Zhuang1
1Department of Mathematics and Information Technology, The Education University of Hong Kong, Hong Kong SAR, P. R. China.
International Journal of Neural Systems
|February 6, 2024
Summary
This study introduces an automated method to evaluate vision-language benchmark datasets. Findings reveal significant quality variations in ground-truth descriptions, with some being unreliable for AI model training.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Large-scale benchmark datasets are vital for advancing AI model development and performance evaluation.
- Existing vision-language datasets like Flickr30k, COCO, and NoCaps pair images with textual descriptions.
Purpose of the Study:
- To propose an automatic method for assessing the quality of large-scale benchmark datasets for vision-language tasks.
- To identify potential issues with the reliability of ground-truth descriptions in these datasets.
Main Methods:
- Development of a novel cross-modal matching model to automatically score textual descriptions against visual images.
- Application of the developed model to evaluate existing vision-language datasets by scoring each image-description pair.
Main Results:
- The automated scoring method shows good agreement with manual evaluations.
- Significant disparities in the quality of ground-truth descriptions across benchmark datasets were identified.
- A notable portion of descriptions were found to be unsuitable as reliable ground-truth references.
Conclusions:
- The proposed automated method effectively assesses vision-language dataset quality.
- There is a critical need for careful scrutiny and utilization of publicly available benchmark datasets.
- Improving dataset quality is essential for the reliable progress of AI systems in vision-language research.
Keywords:
Benchmark datasetsautomated scoringcross-modal deep learningquality evaluationvision-language tasks
