Related Experiment Video
Updated: Jun 28, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Unsupervised Sentence Representation Learning with Frequency-induced Adversarial tuning and Incomplete sentence
Bing Wang1, Ximing Li1, Zhiyao Yang1
1College of Computer Science and Technology, Jilin University, China; Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University, China.
Pre-trained Language Models (PLMs) create biased sentence embeddings due to word frequency. Our Slt-fai framework improves unsupervised sentence representation learning by making embeddings frequency-invariant and emphasizing informative low-frequency words.
Area of Science:
- Natural Language Processing
- Machine Learning
Background:
- Pre-trained Language Models (PLMs) are foundational for Unsupervised Sentence Representation Learning (USRL).
- PLMs exhibit anisotropic embedding spaces due to word frequency sensitivity, causing similarity and information bias, which degrades sentence embedding quality.
Purpose of the Study:
- To address the limitations of PLMs in USRL.
- To propose a novel framework, Slt-fai, for improving sentence representation learning by mitigating frequency-induced biases.
Main Methods:
- Calculated word frequencies from PLM pre-training corpora and assigned frequency labels.
- Developed a similarity discriminator for adversarial tuning to create frequency-invariant embeddings.
- Introduced an incomplete sentence detection task with an information discriminator to emphasize informative low-frequency words.
Main Results:
- Achieved a uniformly frequency-invariant embedding space.
- Enhanced the emphasis on informative low-frequency words.
- Demonstrated Slt-fai's superiority over existing USRL baselines across various backbones and datasets.
Conclusions:
- Slt-fai effectively overcomes frequency-induced biases in PLMs for USRL.
- The framework is flexible, plug-and-play, and enhances sentence embedding quality.
- Slt-fai offers a significant improvement for unsupervised sentence representation learning.
More Related Videos
05:54Eye-tracking to Distinguish Comprehension-based and Oculomotor-based Regressive Eye Movements During Reading
Published on: October 18, 2018
12:49Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
Related Concept Videos
Frequency-dependent Selection
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Associative Learning
Classical conditioning, also known...
Determination of Expected Frequency
Construction of Frequency Distribution
First, make a table with two columns—one with the title of the data that needs to be organized, and the other column for frequency. [Draw a third column for tally marks if needed]. Then, take a look at the items given in the data set and decide if an ungrouped frequency distribution table or a grouped frequency distribution table would be more suitable. If there are large sets of different values, then it is...
Aliasing
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...