Related Experiment Video
Updated: Jul 26, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Dataset construction method of cross-lingual summarization based on filtering and text augmentation
Hangyu Pan1, Yaoyi Xi1, Ling Wang1
1State Key Laboratory of Mathematical Engineering and Advanced Computing, Zhengzhou, China.
Abstract:
Existing cross-lingual summarization (CLS) datasets consist of inconsistent sample quality and low scale. To address these problems, we propose a method that jointly supervises quality and scale to build CLS datasets. In terms of quality supervision, the method adopts a multi-strategy filtering algorithm to remove low-quality samples of monolingual summarization (MS) from the perspectives of character and semantics, thereby improving the quality of the MS dataset. In terms of scale supervision, the method adopts a text augmentation algorithm based on the pretrained model to increase the size of CLS datasets with quality assurance. This method was used to build an English-Chinese CLS dataset and evaluate it with a reasonable data quality evaluation framework. The evaluation results show that the dataset is of good quality and large size. These outcomes show that the proposed method may comprehensively improve quality and scale, thereby resulting in a high-quality and large-scale CLS dataset at a lower cost.
Related Concept Videos
Improving Translational Accuracy
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Cross-Sectional Research
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Random Sampling Method
Data Collection by Survey

