Related Experiment Video
Updated: Nov 22, 2025

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
Single- and Cross-Modality Near Duplicate Image Pairs Detection via Spatial Transformer Comparing CNN
Yi Zhang1, Shizhou Zhang1, Ying Li1,2
1School of Computer Science, National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, Shaanxi Provincial Key Laboratory of Speech & Image Information Processing, Northwestern Polytechnical University, Xi'an 710129, China.
This study introduces a novel spatial transformer comparing convolutional neural network (CNN) for near-duplicate image detection. The model effectively utilizes correlations between image pairs, outperforming existing methods in both single and cross-modality tasks.
Area of Science:
- Computer Vision
- Pattern Recognition
- Machine Learning
Background:
- Near-duplicate image detection is crucial for tasks like content-based image retrieval and copyright protection.
- Current deep learning methods often overlook inter-image correlations, limiting performance.
Purpose of the Study:
- To propose a novel deep learning model for enhanced near-duplicate image pair detection.
- To improve the utilization of correlation information between image pairs.
- To address local deformations in image pairs.
Main Methods:
- A comparing convolutional neural network (CNN) framework with a cross-stream to learn inter-image correlations.
- Integration of a spatial transformer module to handle local image deformations (cropping, scaling, etc.).
- Extensive experiments on benchmark datasets for single-modality and cross-modality (Optical-InfraRed) detection.
Main Results:
- The proposed spatial transformer comparing CNN model achieved superior performance.
- Demonstrated effectiveness on both single-modality and cross-modality near-duplicate detection.
- Outperformed several state-of-the-art methods on popular benchmark datasets.
Conclusions:
- The developed model effectively leverages correlations between image pairs for improved detection.
- The spatial transformer component enhances robustness to local deformations.
- The method shows significant potential for real-world near-duplicate image detection applications.
