Related Experiment Videos
Noisy text categorization
1IDIAP Research Institute, Rue du Simplon 4, 1920 Martigny, Switzerland. vincia@idiap.ch
IEEE Transactions on Pattern Analysis and Machine Intelligence
|December 17, 2005
Summary
Text categorization systems maintain acceptable performance on noisy texts, even with up to 50% word error rates from speech recognition or optical character recognition. Performance loss is manageable for recall up to 70%.
Area of Science:
- Natural Language Processing
- Information Retrieval
- Machine Learning
Background:
- Real-world text data often originates from non-digital sources, introducing noise through extraction processes like speech recognition and optical character recognition.
- The impact of such text noise on the performance of text categorization systems is a critical concern for practical applications.
Purpose of the Study:
- To evaluate the performance degradation of text categorization systems when applied to noisy text data.
- To compare categorization performance on clean versus noisy text versions across different noise levels and sources.
- To propose new metrics for assessing extraction process performance and explaining categorization outcomes.
Main Methods:
- Experiments were conducted using text categorization on both clean and synthetically generated noisy document versions.
- Noisy texts were created using handwriting recognition and simulated optical character recognition, achieving Word Error Rates (WER) from approximately 10% to 50%.
- Performance metrics, including Recall, were analyzed to quantify the impact of noise.
Main Results:
- Text categorization performance shows a noticeable but acceptable decline with increasing noise levels.
- Recall values remain acceptable (up to 60-70%) depending on the specific noise source and its characteristics.
- The study demonstrates that categorization is feasible even with significant text corruption.
Conclusions:
- Text categorization systems can be robust to noise introduced by common extraction methods.
- The proposed performance measures offer a clearer understanding of how extraction errors affect categorization.
- Findings suggest practical viability for categorization tasks involving transcribed speech or OCR-generated text.