Related Experiment Videos
SRM-CSR: unsupervised aspect category detection based on semantic-aware relevance modeling and contextual sentence
Yao Xu1,2, Xian Mu1,3, Ketong Liu1
1School of Computer Science and Engineering, Macau University of Science and Technology, Macau, 999078, China.
Scientific Reports
|June 1, 2026
Summary
This study introduces a new framework for unsupervised aspect category detection, improving pseudo-label generation by combining lexical and sentence-level information. The novel approach enhances model performance in identifying underlying aspect categories without labeled data.
Area of Science:
- Natural Language Processing
- Machine Learning
- Artificial Intelligence
Background:
- Unsupervised aspect category detection identifies implicit topics in text without labeled data.
- Current methods often generate suboptimal pseudo-labels, hindering model training and performance.
- Existing approaches struggle with aspect discriminability and rely on general pre-trained models.
Purpose of the Study:
- To propose a novel framework (SRM-CSR) for high-quality pseudo-label generation in unsupervised aspect category detection.
- To integrate lexical-level aspect-relevant information and sentence-level contextual representations.
- To improve the accuracy and effectiveness of aspect category detection models.
Main Methods:
- Extracting Aspect-Relevant Terms (ARTs) based on domain specificity and semantic stability.
- Employing an entropy-driven mechanism to weight aspect terms within sentences.
- Utilizing Sentence-BERT for contextual sentence encoding and similarity computation.
- Generating pseudo-labels independently at lexical and sentence levels, then consolidating consistent results.
- Training a neural classifier using a post-trained Domain Knowledge BERT.
Main Results:
- The proposed SRM-CSR framework effectively generates high-quality pseudo-labels.
- The integrated approach significantly improves unsupervised aspect category detection.
- Achieved an average improvement of 2.9 percentage points in macro-F1 over baseline methods on three real-world datasets.
Conclusions:
- The novel pseudo-labeling strategy in SRM-CSR is effective for unsupervised aspect category detection.
- Combining lexical and sentence-level information enhances aspect term identification and model performance.
- The framework offers a robust solution for improving models trained without annotated data.