Related Experiment Video
Updated: Sep 11, 2025

05:49
Author Spotlight: Analgesic Effect of Tuina on Rat Models with Compression of the Dorsal Root Ganglion Pain
Published on: July 14, 2023
1.5K
LRQuant+: A Unified and Learnable Framework to Post-Training Quantization for Transformer-Based Large Foundation
IEEE Transactions on Pattern Analysis and Machine Intelligence
|August 14, 2025
Summary
This study introduces LRQuant, a novel post-training quantization method for large foundation models. LRQuant optimizes scaling factors and uses a new loss function to improve model efficiency and accuracy across diverse scenarios.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Computer Vision
Background:
- Post-training quantization (PTQ) accelerates inference and reduces memory for large foundation models (LFMs) without retraining.
- Existing PTQ methods struggle with hand-crafted scaling factors, ignore directional quantization errors, and lack broad applicability.
- Current quantization error metrics (e.g., L2-norm) do not capture directional shifts, leading to suboptimal performance.
Purpose of the Study:
- To develop a unified, learnable, and robust post-training quantization framework (LRQuant) for transformer-based LFMs.
- To address limitations of existing PTQ methods by introducing learnable scaling factors and a novel loss function.
- To provide a comprehensive evaluation across diverse LFMs and quantization scenarios, including challenging low-bit settings.
Main Methods:
- Introduced a block-wise learnable paradigm for optimal scaling factor determination, initialized with logarithmic activation equivalents.
- Proposed a novel Negative Logarithm of Cosine Similarity (NLC) loss to better capture quantization errors beyond MSE.
- Developed LRQuant+ with a dynamic loss weighting scheme, learnable rotation vectors, and a two-branch optimization for error propagation and reconstruction.
Main Results:
- LRQuant and LRQuant+ demonstrate superior performance across various LFMs, including LLMs, ViTS, and MLLMs.
- The methods achieve effectiveness in both weight-activation and weight-only quantization, particularly in challenging W4A4 and W2A16 scenarios.
- Experimental results validate the unified applicability and robustness of the proposed LRQuant framework.
Conclusions:
- LRQuant offers a significant advancement in post-training quantization for LFMs, improving efficiency and accuracy.
- The learnable approach and novel NLC loss effectively mitigate quantization errors and enhance model robustness.
- The framework's versatility across different models and quantization bit-widths makes it a valuable tool for deploying LFMs.
Related Concept Videos
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Estimation of the Physical Quantities
5.9K
On many occasions, physicists, other scientists, and engineers need to make estimates of a particular quantity. These are sometimes referred to as guesstimates, order-of-magnitude approximations, back-of-the-envelope calculations, or Fermi calculations. The physicist Enrico Fermi was famous for his ability to estimate various kinds of data with surprising precision. Estimating does not mean guessing a number or a formula at random. Instead, estimation means using prior experience and sound...
5.9K
Transformers with Off-Nominal Turns Ratios
205
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
205
The Ideal Transformer
890
In single-phase two-winding transformers, two windings are coiled around a magnetic core characterized by cross-sectional area A and magnetic permeability μ. A phasor current i1 enters the left winding while i2 exits the right winding, establishing the fundamental working of the transformer through electromagnetic principles.
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
890
Energy Losses in Transformers
974
In an ideal transformer, it is assumed that there are no energy losses, and, hence, all the power at the primary winding is transferred to the secondary winding. However, in reality, the transformers always have some energy losses, and, hence, the output power obtained at the secondary winding is less than the input power at the primary winding due to energy losses.
There are four main reasons for energy losses in transformers.
The first cause can be the high resistance of the...
There are four main reasons for energy losses in transformers.
The first cause can be the high resistance of the...
974
Linear Approximation in Frequency Domain
131
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
131

