CLTD-LP: an optimized top-down clustering approach with linear prefix trees for scalable frequent pattern discovery
M Sinthuja1, M Diviya2, P Saranya2
1School of Computer Science and Engineering, Vellore Institute of Technology, Vellore, India. sinthuja.m@vit.ac.in.
None:
The extraction of frequent itemsets and association rules is a fundamental challenge in data mining and holds significant importance within the field. Mining techniques utilising Linear Prefix (LP) growth association rules employ a bottom-up methodology that necessitates a conditional pattern base and a conditional LP-tree for the extraction of frequent itemsets. This research proposes a Linear Prefix tree utilising a Top-Down Approach with Clustering (CLTDLP) method, to address the limitations of the current LP-growth algorithm. The suggested CLTD-LP algorithm employs a top-down methodology; it generates a subheader table to effectively mine common items, rendering the CLTD-LP algorithm more advantageous as it does not create a conditional pattern base or LP-tree. The proposed methodology enhances the algorithm's efficiency regarding execution time and memory utilisation. Overall, across three benchmark datasets, the proposed CLTD-LP algorithm consistently achieves better average reduction in terms of runtime and memory than the existing LP-growth, Ordered Frequent Itemsets Matrix (OFIM) and Simple and Scalable Frequent Itemset Mining(SSFIM) algorithms.
More Related Videos
06:01Visualization and Quantification of High-Dimensional Cytometry Data using Cytofast and the Upstream Clustering Methods FlowSOM and Cytosplore
Published on: December 12, 2019
07:28JUMPn: A Streamlined Application for Protein Co-Expression Clustering and Network Analysis in Proteomics
Published on: October 19, 2021
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
