Compression of Deep Neural Networks based on quantized tensor decomposition to implement on reconfigurable hardware

Amirreza Nekooei1, Saeed Safari1

  • 1University of Tehran, Iran.

Summary

This study introduces tensor decomposition to compress deep neural network (DNN) parameters, significantly reducing memory usage for mobile AI applications. The method maintains network accuracy while enabling efficient deployment on hardware like FPGAs.