Related Experiment Video
Updated: Jan 16, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
IFE-CMT: Instance-Aware Fine-Grained Feature Enhancement Cross Modal Transformer for 3D Object Detection.
Xiaona Song1, Haozhe Zhang1, Haichao Liu1
1School of Mechanical Engineering, North China University of Water Resources and Electric Power, Zhengzhou 450045, China.
This study introduces the Instance-aware Fine-grained feature Enhancement Cross Modal Transformer (IFE-CMT) model to improve multi-modal 3D object detection. The IFE-CMT model significantly enhances the detection accuracy of small objects, achieving better performance on benchmark datasets.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Multi-modal 3D object detection algorithms have advanced significantly.
- Current methods often overlook fine-grained features, impacting small object detection accuracy.
- Existing fusion strategies can lead to a decline in performance for detecting smaller objects.
Purpose of the Study:
- To propose a novel model, the Instance-aware Fine-grained feature Enhancement Cross Modal Transformer (IFE-CMT), for improved multi-modal 3D object detection.
- To enhance the detection accuracy of small objects within complex scenes.
- To address the limitations of current fusion strategies in capturing fine-grained object representations.
Main Methods:
- Developed an Instance feature Enhancement Module (IE-Module) for accurate multi-modal feature extraction and enhancement.
- Introduced a new point cloud branch network to expand the receptive field and improve semantic expression.
- Designed a cross-modal transformer architecture focusing on fine-grained feature enhancement.
Main Results:
- The IFE-CMT model demonstrated improved performance on the nuScenes dataset compared to the CMT model.
- Achieved a 2.1% and 1.9% increase in mAP on the validation and test sets, respectively.
- Significantly improved mAP for small objects like bicycles (6.6%) and motorcycles (3.7%).
Conclusions:
- The proposed IFE-CMT model effectively enhances fine-grained feature representations for multi-modal 3D object detection.
- IFE-CMT shows superior performance, particularly in detecting small objects, outperforming existing methods.
- The model offers a promising approach to address the challenges in small object detection within autonomous driving systems.
Related Concept Videos
Transformers
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
Transformation
Types Of Transformers
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
Transformers with Off-Nominal Turns Ratios
Improving Translational Accuracy
Improving Translational Accuracy
