Related Experiment Video
Updated: Jun 10, 2026

14:27
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
A New-Fangled FES-k-Means Clustering Algorithm for Disease Discovery and Visual Analytics
1GIS Research Laboratory for Geographic Medicine, Advanced Geospatial Analysis Laboratory, Department of Geography & Environmental Resources, Southern Illinois University, 1000 Faner Drive, MC 4514, Carbondale, IL 62901-4514, USA.
EURASIP Journal on Bioinformatics & Systems Biology
|August 7, 2010
Summary
A new Fast, Efficient, and Scalable k-means (FES-k-means) algorithm improves clustering efficiency and speed. This method enhances data mining and geospatial analysis, potentially linking environmental factors to disease mechanisms.
Area of Science:
- Data Science
- Computational Biology
- Geospatial Analysis
Background:
- The original k-means clustering technique can be computationally intensive for large datasets.
- Efficient data clustering is crucial for knowledge discovery and data mining applications.
- Geospatial data analysis has significant implications for understanding disease patterns and mechanisms.
Purpose of the Study:
- To evaluate the performance and quality of a new algorithm, the Fast, Efficient, and Scalable k-means (FES-k-means) algorithm.
- To provide evidence for FES-k-means as an improvement over the original k-means clustering technique.
- To explore the application of FES-k-means in analyzing large geospatial datasets for disease mechanism discovery.
Main Methods:
- Developed the FES-k-means algorithm, a hybrid approach combining k-d tree data structure, the original k-means algorithm, and Mashor's adaptation rate.
- Tested the algorithm on two real-world datasets and one synthetic dataset.
- Applied the algorithm in a two-step process: first on data trained by the MIL-SOM method, then on untrained data.
Main Results:
- The FES-k-means algorithm produced clusters comparable to the original k-means method.
- Runtime comparisons demonstrated significantly faster performance for FES-k-means.
- The algorithm efficiently analyzed large geospatial datasets, showing potential for disease mechanism discovery.
Conclusions:
- FES-k-means offers a substantial improvement in speed and efficiency for k-means clustering.
- The two-step training and clustering approach enhances knowledge discovery capabilities.
- The algorithm's efficiency in geospatial analysis may aid in discovering links between environmental exposures, such as water service lines, and disease patterns, like elevated blood lead levels in Chicago.
