提高医疗数据集的分类准确度,使用基于混合距离和集群改进的K-means集群方法
Hussein A A Al-Khamees1, Mudatheer M Al-Slivani2, Mayameen S Kadhim3
1Computer Techniques Engineering Department, College of Engineering and Technology, Al-Mustaqbal University, 51001, Babil, Iraq. Hussein.Alkhamees@uomus.edu.iq.
Scientific reports
|January 25, 2026
概括
这项研究通过引入混合距离度量和精细化步骤来增强K-Means集群用于医学数据分析,显著提高准确性和集群质量,以便更好地做出临床决策.
科学领域:
- 机器学习是机器学习.
- 数据科学是数据科学.
- 医疗信息学医学信息学
背景情况:
- 经典的K-Means集群与医疗数据扎,原因是距离指标不足,缺少分配后的精细化.
- 这些局限性导致集群凝聚力较差,医疗数据集中的错误分组.
- 现有的方法往往无法完全捕捉复杂的医疗数据结构.
研究的目的:
- 为医疗数据分析提出一种新的增强的K-Means集群框架.
- 通过结合混合距离度量和集群精细化机制来解决K-Means的局限性.
- 提高医疗保健应用中的聚类的准确性,可解释性和稳定性.
主要方法:
- 开发了一种混合距离方法,结合了可调节权重的共弦值和城市区块 (曼哈顿) 度量.
- 实施了一个集群精细化机制,使用Z-score异常值检测来重新分配远距离的样本.
- 使用多个性能指标评估了威斯康星州乳腺癌 (BCW) 和心脏病数据集的框架.
主要成果:
- 增强的K-Means实现了高精度:BCW为0.9825和心脏病为0.9000.
- 显著优于传统的欧几里德和基于等号的K-Means方法.
- 在同质性得分方面表现出显著的改善,表明集群质量和分离性得到了增强.
结论:
- 拟议的混合K-Means框架为医疗数据集群提供了实用和有效的增强.
- 混合距离度量和精细化步骤的结合导致比现有方法更优越的性能.
- 这种方法在医疗数据分析和临床决策中改善无监督学习方面具有显著的潜力.
相关概念视频
Cluster Sampling Method
14.3K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
14.3K
Vesicular Tubular Clusters
3.1K
After budding out from the ER membrane, some COPII vesicles lose their coat and fuse with one another to form larger vesicles and interconnected tubules called vesicular tubular clusters or VTCs. These clusters constitute a compartment at the ER-Golgi interface known as ERGIC (Endoplasmic Reticulum Golgi Intermediate Compartment). The ERGIC is a mobile membrane-bound cargo transport system that sorts proteins secreted from ER and delivers them to the Golgi.
With the help of motor proteins such...
With the help of motor proteins such...
3.1K
Chromatographic Methods: Classification
3.8K
Chromatographic techniques are classified in three ways: the classification is based on the physical state of the stationary and mobile phases, how the mobile phase and the stationary phase contact each other, or through the chemical or physical processes that isolate the components of the sample. Typically, the mobile phase is either a liquid or gas, while the stationary phase is either a solid or a liquid layer applied to a solid surface.
Chromatographic techniques are typically named by...
Chromatographic techniques are typically named by...
3.8K
Methods of Classification and Identification
1.1K
Bacterial identification relies on a diverse array of techniques to classify and understand microorganisms, each tailored to uncover specific characteristics. Traditional morphological approaches, while still valuable, are limited for closely related or structurally simple organisms. Modern methods integrate biochemical, serological, genetic, and advanced molecular tools to achieve greater accuracy.Morphological and Biochemical TechniquesMorphological characteristics, such as cell shape and...
1.1K
Distance Problem
62
When an object's velocity changes over time, the total distance traveled can be determined by summing small displacement intervals over short increments. This approach approximates the true distance through numerical summation and the use of integral calculus. An estimate of the total displacement can be obtained by measuring velocity at regular intervals and multiplying each value by the corresponding time step.If a runner accelerates over the first three seconds of a race, speed measurements...
62
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K


