Related Experiment Video
Updated: Jul 13, 2025

Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese
Published on: April 1, 2016
Research on performance variations of classifiers with the influence of pre-processing methods for Chinese short text
Dezheng Zhang1,2, Jing Li1,2, Yonghong Xie1,2
1School of Computer and Communication Engineering, University of Science and Technology Beijing, Haidian, Beijing, China.
Preprocessing Chinese text, including word segmentation and stop word removal, significantly boosts text classification performance. This study demonstrates the positive impact of systematic preprocessing on machine and deep learning models for Chinese short text.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Machine Learning
Background:
- Text preprocessing is crucial for Chinese text classification.
- Existing research primarily focuses on English text preprocessing.
- Limited studies explore preprocessing's impact on diverse Chinese text classification algorithms.
Purpose of the Study:
- To experimentally compare Chinese text preprocessing methods.
- To evaluate their influence on fifteen common text classifiers.
- To analyze performance under various conditions like evaluation metrics and classifier types.
Main Methods:
- Utilized three standard Chinese preprocessing techniques: word segmentation, stop word removal, and symbol removal.
- Tested fifteen classifiers on two Chinese datasets.
- Analyzed classification results using metrics like macro-F1, considering different preprocessing combinations and classifier selections.
Main Results:
- Most classifiers showed improved performance after applying appropriate preprocessing.
- Systematic preprocessing positively impacted Chinese short text classification.
- Achieved macro-F1 scores of 92.13% and 91.99%, outperforming baselines by 0.3% and 2% respectively.
Conclusions:
- Systematic application of preprocessing methods enhances Chinese short text classification.
- Word segmentation, stop word removal, and symbol removal are effective preprocessing steps.
- Preprocessing benefits both machine learning and deep learning models.
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Survival Tree
Building a Survival Tree
Constructing a...
Classification of Systems-II
Classification of Leukocytes
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...

