An Online Weighted Bayesian Fuzzy Clustering Method for Large Medical Data Sets.
Cong Zhang1, Jing Xue2, Xiaoqing Gu3
1School of Computer and Information Engineering, Nantong Polytechnic College, Nantong 226001, Jiangsu, China.
This study introduces an online clustering framework to efficiently process large medical datasets. The method accelerates clustering and reduces memory usage, outperforming previous techniques in speed and performance.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Data Science
Background:
- The proliferation of AI and wearable devices generates massive medical datasets exceeding single-machine memory capacity.
- Processing large-scale medical data demands significant computational resources, increasing time and hardware requirements.
- Existing data processing methods struggle with the scale and complexity of modern health informatics.
Purpose of the Study:
- To develop an efficient online clustering framework for large-scale medical datasets.
- To reduce memory consumption and accelerate data processing for health informatics.
- To improve clustering performance on diverse medical data types.
Main Methods:
- A novel online clustering framework is proposed, processing data in blocks.
- Weighting clustering is applied to each data block to derive cluster centers and weights.
- Final cluster centers are computed by aggregating block-specific centers and weights.
Main Results:
- The proposed method significantly reduces processing time compared to traditional approaches.
- Experimental results demonstrate superior clustering performance across standard, cancer, and CT image datasets.
- The framework effectively handles large datasets that cannot be loaded into memory at once.
Conclusions:
- The online clustering framework offers an efficient solution for processing large medical datasets.
- This approach accelerates clustering and minimizes memory footprint without compromising performance.
- The method shows promise for applications in health informatics and medical data analysis.
More Related Videos
06:01Visualization and Quantification of High-Dimensional Cytometry Data using Cytofast and the Upstream Clustering Methods FlowSOM and Cytosplore
Published on: December 12, 2019
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Statistical Software for Data Analysis and Clinical Trials
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Biostatistics: Overview
Discrete variables are...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
