Related Experiment Video
Updated: Oct 3, 2025

Large Scale Energy Efficient Sensor Network Routing Using a Quantum Processor Unit
Published on: September 8, 2023
Optimal Architecture of Floating-Point Arithmetic for Neural Network Training Processors.
Muhammad Junaid1, Saad Arslan2, TaeGeon Lee1
1Department of Electronics, College of Electrical and Computer Engineering, Chungbuk National University, Cheongju 28644, Korea.
This study introduces an optimized mixed-precision accelerator for Artificial Intelligence of Things (AIoT) devices, enabling efficient on-device training and inference. The new design significantly reduces size and energy consumption while maintaining high accuracy for edge AI applications.
Area of Science:
- Computer Engineering
- Artificial Intelligence
- Hardware Acceleration
Background:
- Artificial Intelligence (AI) and the Internet of Things (AIoT) are key drivers of the fourth industrial revolution.
- Current AIoT research focuses on inference accelerators, but training capabilities are increasingly needed for self-supervised and semi-supervised learning.
- High-precision floating-point operations for training demand significant area and energy, posing challenges for edge devices.
Purpose of the Study:
- To develop an energy-efficient and compact accelerator for AIoT devices capable of both inference and training.
- To investigate optimal floating-point formats (32, 24, 16 bits, and mixed precision) for low-power, small-sized edge applications.
- To achieve high accuracy in neural network training and inference on edge devices.
Main Methods:
- Proposed a novel accelerator architecture incorporating mixed-precision floating-point training (32, 24, 16 bits).
- Verified accelerator performance on FPGA for inference and training using the MNIST dataset.
- Implemented the optimized mixed-precision accelerator on ASIC (TSMC 65nm) for area and energy analysis.
Main Results:
- Achieved over 93% accuracy using a combination of 24-bit custom FP and 16-bit Brain FP formats.
- ASIC implementation showed an active area of 1.036 × 1.036 mm² and 4.445 µJ energy consumption per image training.
- Reduced size by 4.7 times and energy consumption by 3.91 times compared to 32-bit architectures.
Conclusions:
- The optimized mixed-precision accelerator significantly reduces area and energy consumption for AIoT edge devices.
- This architecture supports high-accuracy training and inference, crucial for advancing AIoT applications.
- The proposed CNN structure with an optimized data path is vital for developing compact, low-power, high-accuracy AIoT solutions.
Related Concept Videos
Ampere-Maxwell's Law: Problem-Solving
To solve the problem, we can use the equations from the analysis of an RC circuit and Maxwell's version of Ampère's law.
For the first part of...
Parallel Processing
Numerical Calculations
The solution to a problem is obtained using different methods. While manually solving algebraic symbols is one of the most common methods, the graphical method is often preferred. Computers...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Fast Fourier Transform
The computational efficiency of the FFT becomes...
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...

