Related Experiment Video
Updated: Dec 22, 2025

PIPEMAT-RS: Development and Validation of a Standardized MATLAB Pipeline for Resting-State EEG Preprocessing
Published on: June 6, 2025
The influence of preprocessing on text classification using a bag-of-words representation
Yaakov HaCohen-Kerner1, Daniel Miller1, Yair Yigal1
1Dept. of Computer Science, Jerusalem College of Technology - Lev Academic Center, Jerusalem, Israel.
Exploring text classification preprocessing methods is crucial for improving accuracy. Systematic experiments show that combining preprocessing techniques, like stopword removal and lowercasing, consistently enhances text classification performance.
Area of Science:
- Natural Language Processing
- Machine Learning
Background:
- Text classification (TC) is vital for organizing information.
- Preprocessing is a common step in TC applications.
- The impact of various preprocessing combinations on TC performance is not fully understood.
Purpose of the Study:
- To systematically investigate the effect of different preprocessing method combinations on text classification accuracy.
- To identify optimal preprocessing strategies for benchmark text corpora.
Main Methods:
- Conducted extensive text classification experiments using four benchmark corpora.
- Evaluated all possible combinations of five to six basic preprocessing methods.
- Utilized three machine learning methods with standard training and testing protocols.
- Employed a bag-of-words (BOW) representation for text data.
Main Results:
- At least one preprocessing combination significantly improved TC accuracy across all tested datasets.
- Stopword removal was the most effective single method for three datasets.
- For one dataset, lowercasing was the only effective single method, with spelling correction and lowercasing yielding the best results.
- Minimal improvements were observed with HTML tag removal, spelling correction, punctuation removal, and character reduction in some cases.
Conclusions:
- Systematic exploration of preprocessing method combinations is highly recommended for improving text classification.
- The optimal preprocessing strategy varies depending on the dataset.
- Preprocessing significantly enhances TC accuracy, especially when using BOW representations.
More Related Videos
06:09P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation
Published on: September 8, 2023
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
pre-mRNA Processing
Once about 20-40 ribonucleotides have been joined together by RNA polymerase, a group of enzymes adds a “cap” to the 5’ end of the growing transcript. In this process, a 5’ phosphate is replaced by modified guanosine that has a methyl group attached to it (7-Methyl...
Pre-mRNA Processing: Modification of pre-mRNA Ends
Once about 20-40 ribonucleotides have been joined together by RNA polymerase, a group of enzymes adds a cap to the 5' end of the growing transcript. In this process, a 5' phosphate is replaced by modified guanosine that has a methyl group attached (7-methyl guanosine). This 5' cap helps...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Classification of Systems-II