TransXNet: Learning Both Global and Local Dynamics With a Dual Dynamic Token Mixer for Visual Recognition

Summary

This study introduces a novel dual dynamic token mixer (D-Mixer) for vision networks, enhancing performance by enabling dynamic adaptation to input data. The proposed TransXNet model achieves superior accuracy and efficiency in image classification and dense prediction tasks.