Cluster Validation Method for Determining the Number of Clusters in Categorical Sequences
This study introduces a new cluster validity index (CVI) for evaluating categorical sequence clusters. The proposed method enhances clustering quality by assessing intracluster compactness and intercluster separation.
Area of Science:
- Computational biology
- Machine learning
- Data mining
Background:
- Cluster validation is crucial for machine learning systems.
- Categorical sequences are increasingly common in real-world applications.
- Existing validation methods are insufficient for sequence data.
Purpose of the Study:
- To address the challenge of validating clusters in categorical sequences.
- To propose a novel cluster validity index (CVI) for sequence data.
- To develop a robust clustering algorithm for categorical sequences.
Main Methods:
- A novel CVI combining intracluster structural compactness and intercluster structural separation.
- A partition-based clustering algorithm with deterministic initialization and noise cluster elimination.
- Integration of the algorithm and CVI for determining the optimal number of clusters.
Main Results:
- Demonstrated effectiveness on protein sequences and real-world datasets.
- The proposed CVI accurately measures the quality of sequence clusters.
- The clustering algorithm yields high-quality results.
Conclusions:
- The novel CVI and clustering algorithm offer a robust solution for categorical sequence validation.
- The method effectively determines the number of clusters in sequence datasets.
- This work advances the field of cluster validation for complex sequence data.
More Related Videos
07:59Author Spotlight: Alignment of Synchronized Time-Series Data Using the Characterizing Loss of Cell Cycle Synchrony Model for Cross-Experiment Comparisons
Published on: June 9, 2023
05:12ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Construction of Frequency Distribution
First, make a table with two columns—one with the title of the data that needs to be organized, and the other column for frequency. [Draw a third column for tally marks if needed]. Then, take a look at the items given in the data set and decide if an ungrouped frequency distribution table or a grouped frequency distribution table would be more suitable. If there are large sets of different values, then it is...
Survival Tree
Building a Survival Tree
Constructing a...
Quantifying and Rejecting Outliers: The Grubbs Test
