Related Experiment Video
Updated: May 31, 2025

11:23
Lensless Fluorescent Microscopy on a Chip
Published on: August 17, 2011
17.6K
A Low-Power General Matrix Multiplication Accelerator with Sparse Weight-and-Output Stationary Dataflow
1Research Center for Novel Computing Sensing and Intelligent Processing, Zhejiang Lab, Hangzhou 311100, China.
Micromachines
|January 25, 2025
Summary
This study presents a novel sparse General Matrix Multiplication (GEMM) accelerator for efficient machine learning on resource-constrained devices. The approach enhances computing and energy efficiency by optimizing data movement and buffer utilization.
Area of Science:
- Computer Science
- Machine Learning Hardware
Background:
- General Matrix Multiplication (GEMM) is computationally intensive, limiting its use in resource-constrained environments.
- Existing data reuse methods for GEMM do not fully exploit potential reductions in data movement.
Purpose of the Study:
- To develop a sparse GEMM accelerator that improves efficiency for machine learning on edge devices.
- To address challenges in storing and processing sparse matrices in hardware.
Main Methods:
- Introduced a weight-and-output stationary (WOS) dataflow and distributed buffer architecture for sparse GEMM.
- Developed an adaptable mapping scheme for compressed GEMM and an offline sparsity-aware shuffle strategy for weights.
- Implemented a low-cost sparse computing method with globally shared inputs for high throughput.
Main Results:
- The sparse GEMM accelerator processes matrices in a compressed format, eliminating on-chip weight and partial sum transfers.
- The sparsity-aware shuffle strategy balances buffer utilization and minimizes waste for irregular weight matrices.
- FPGA experiments demonstrated 1.73x better computing efficiency and 1.36x higher energy efficiency compared to existing methods.
Conclusions:
- The proposed sparse GEMM accelerator effectively reduces data movement and improves hardware efficiency for machine learning.
- The novel dataflow, buffer architecture, and weight management strategy enable efficient processing of sparse matrices.
- This work offers a viable solution for deploying computationally demanding machine learning models on resource-limited platforms.
Related Concept Videos
Parallel Processing
144
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
144
Acceleration Vectors
7.9K
In everyday conversation, accelerating means speeding up. Acceleration is a vector in the same direction as the change in velocity, Δv, therefore the greater the acceleration, the greater the change in velocity over a given time. Since velocity is a vector, it can change in magnitude, direction, or both. Thus acceleration is a change in speed or direction, or both. For example, if a runner traveling at 10 km/h due east slows to a stop, reverses direction, and continues their run at 10 km/h...
7.9K
Ampere-Maxwell's Law: Problem-Solving
529
A parallel-plate capacitor with capacitance C, whose plates have area A and separation distance d, is connected to a resistor R and a battery of voltage V. The current starts to flow at t = 0. What is the displacement current between the capacitor plates at time t? From the properties of the capacitor, what is the corresponding real current?
To solve the problem, we can use the equations from the analysis of an RC circuit and Maxwell's version of Ampère's law.
For the first part of...
To solve the problem, we can use the equations from the analysis of an RC circuit and Maxwell's version of Ampère's law.
For the first part of...
529
Vector Operations
1.1K
Vectors are physical quantities that have both magnitude and direction. The vector operations include addition, subtraction, and scalar multiplication.
A vector multiplied by a scalar value is called scalar multiplication. The result obtained is a new vector with a different magnitude. If the scalar is positive, the direction of the vector remains the same, but if it is negative, the direction of the vector is reversed. For example, the product of the mass and velocity yields the momentum.
A vector multiplied by a scalar value is called scalar multiplication. The result obtained is a new vector with a different magnitude. If the scalar is positive, the direction of the vector remains the same, but if it is negative, the direction of the vector is reversed. For example, the product of the mass and velocity yields the momentum.
1.1K
Distributed Loads
508
Distributed loads are a common type of load that engineers and scientists encounter in various practical situations. Distributed loads often refer to a type of load spread over a surface or a structure and can be modeled as continuous force per unit area.
For example, consider a bookshelf filled with books stacked vertically adjacent to each other. The weight of the books is evenly distributed over the length of the shelf. As a result, the pressure at different locations on the surface of the...
For example, consider a bookshelf filled with books stacked vertically adjacent to each other. The weight of the books is evenly distributed over the length of the shelf. As a result, the pressure at different locations on the surface of the...
508
Fast Decoupled and DC Powerflow
164
The fast decoupled power flow method addresses contingencies in power system operations, such as generator outages or transmission line failures. This method provides quick power flow solutions, essential for real-time system adjustments. Fast decoupled power flow algorithms simplify the Jacobian matrix by neglecting certain elements, leading to two sets of decoupled equations:
164

