Related Experiment Video
Updated: Mar 14, 2026

Decoding Natural Behavior from Neuroethological Embedding
Published on: October 3, 2025
Improving the text classification using clustering and a novel HMM to reduce the dimensionality.
A Seara Vieira1, L Borrajo1, E L Iglesias1
1Department of Computer Science, Higher Technical School of Computer Engineering, University of Vigo, 32004 Ourense, Spain.
This study introduces a new method for text classification by reducing document dimensionality using clustering and Hidden Markov Models (HMMs). This approach improves computational efficiency and classification performance on benchmark datasets.
Area of Science:
- Computer Science
- Machine Learning
- Natural Language Processing
Background:
- High dimensionality in text representations burdens computational processes in machine learning.
- Effective document representation is crucial for successful text classification.
- Existing methods require optimization for large-scale real-world data.
Purpose of the Study:
- To propose a novel dimensionality reduction technique for text representation.
- To enhance the efficiency and performance of text classification systems.
- To evaluate the effectiveness of the proposed method against established techniques.
Main Methods:
- Document representation using a clustering technique and a Hidden Markov Model (HMM).
- Dimensionality reduction applied as a preprocessing step for text classification.
- Evaluation using k-Nearest Neighbors (k-NN) and Support Vector Machine (SVM) classifiers on OHSUMED and TREC corpora.
Main Results:
- The proposed dimensionality reduction technique yielded highly satisfactory experimental results.
- Performance was superior compared to commonly used methods like InfoGain.
- Statistical tests confirmed the technique's suitability for text classification preprocessing.
Conclusions:
- The novel clustering and HMM-based dimensionality reduction method is effective for text classification.
- The technique offers a significant improvement in efficiency and performance.
- This approach is well-suited for the preprocessing stage in text classification tasks.
More Related Videos
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Related Concept Videos
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Classification of Systems-II
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...