Related Experiment Video
Updated: May 6, 2026

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
Research on the similarity calculation of short text in the terminology domain based on siamese BERT model
Fei Chen1, Zhenling Zhang1, Yangli Jia2
1School of Computer Science, and the Knowledge Engineering & Terminology Research Centre (KETRC), Liaocheng University, Liaocheng, 252000, Shandong, China.
Abstract:
This study focuses on the task of term-domain short-text similarity computation. It addresses two main challenges: dataset scarcity and insufficient deep semantic extraction. To solve these issues, we first construct the multi-domain Terminology Definition Similarity (TDS) dataset using an automated pipeline. This pipeline combines data generation (based on the GPT-4 model) with a manual verification process involving expert quality control. The design ensures the production of high-quality data. We then present an innovative model named SDQKC (Sbert + dynamic QK + contrast). The model optimizes the Siamese BERT network through a dynamic QK co-attention mechanism and enhances its deep-level semantic understanding by incorporating contrastive learning. Experimental results show that the SDQKC model achieves Pearson correlation coefficients of 0.69384 on the BQ Corpus (Bank Question Corpus) and 0.69511 on the TDS dataset. These results significantly outperform other baseline models, demonstrating the effectiveness and superiority of the proposed methodology.
Related Concept Videos
SBAR I: Understanding the Concept
Standardized methods of communication have been developed to ensure that information is...
Classification of Systems-II
Modeling and Similitude
Self-Evaluation Maintenance Model
Factors Influencing Attraction III: Similarity
Causes of Similarity-Dissimilarity Effect