Related Experiment Video
Updated: Jun 10, 2025

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Text classification algorithm of tourist attractions subcategories with modified TF-IDF and Word2Vec
Lu Xiao1,2,3, Qiaoxing Li1,3, Qian Ma4
1School of Management, Guizhou University, Guiyang, China.
This study introduces a new text classification method for tourist attraction descriptions, outperforming existing models like BERT. The novel approach enhances text representation for better accuracy and stability in complex datasets.
Area of Science:
- Natural Language Processing
- Text Mining
- Machine Learning
Background:
- Text classification is crucial for managing big data, but its application to specialized domains like tourist attractions is underexplored.
- Existing methods struggle with complex, imbalanced datasets common in professional fields.
Purpose of the Study:
- To develop and validate a novel text representation and classification method for tourist attraction descriptions.
- To address the challenges of classifying complex, imbalanced, and domain-specific text data.
Main Methods:
- Constructed a corpus of tourist attraction descriptions using web crawlers.
- Proposed a hybrid text representation combining Word2Vec embeddings with TF-IDF-CRF-POS weighting.
- Integrated the representation with seven common classifiers (DT, SVM, LR, NB, MLP, RF, KNN) for multi-class classification.
Main Results:
- The proposed algorithm achieved higher accuracy (2.29%) and F1-scores (macro-F1: 5.55%, micro-F1: 2.90%) compared to traditional methods and BERT.
- Demonstrated superior performance on imbalanced datasets, identifying rare categories effectively.
- Exhibited enhanced stability across datasets of varying sizes.
Conclusions:
- The novel text representation and classification algorithm offers superior performance and robustness for professional domain text.
- The method is practical and provides a valuable reference for vector expression and classification of complex Chinese text datasets.
More Related Videos
Related Concept Videos
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Classification of Systems-II
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Methods of Classification and Identification
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...

