Related Experiment Videos
A clustering method based on rough sets and its application to knowledge discovery in the medical database
S Hirano1, S Tsumoto, T Okuzaki
1Department of Medical Informatics, Shimane Medical University, School of Medicine, Izumo, Shimane 691-8501, Japan. hirano@ieee.org
Studies in Health Technology and Informatics
|October 18, 2001
Summary
This study introduces a novel clustering method using Rough Sets for mixed data types, enhancing medical knowledge discovery. The approach effectively identifies key diagnostic factors for diseases like meningoencephalitis.
Area of Science:
- Data Mining
- Artificial Intelligence
- Medical Informatics
Background:
- Clustering mixed data types (nominal and numerical) presents challenges in knowledge discovery.
- Existing methods may struggle with the inherent differences between attribute types.
- Effective classification requires robust similarity measures for diverse data.
Purpose of the Study:
- To propose a novel clustering method for nominal and numerical data using Rough Sets.
- To apply this method for knowledge discovery in medical databases.
- To identify primary diagnostic factors from complex medical data.
Main Methods:
- Utilizing Rough Sets theory for clustering based on indiscernibility relations.
- Defining similarity using a combination of Hamming distance (nominal) and Mahalanobis distance (numerical).
- Modifying equivalence relations to prevent the excessive generation of small categories.
Main Results:
- The proposed method successfully clusters data with both nominal and numerical attributes.
- Validation on a meningoencephalitis diagnosis database demonstrated effectiveness.
- The method identified significant factors contributing to disease diagnosis.
Conclusions:
- The Rough Set-based clustering method is effective for mixed-type data in medical knowledge discovery.
- This approach facilitates the identification of critical diagnostic indicators.
- The method offers a robust solution for analyzing complex medical databases.