通过一种新的学习方法,提高K-最近邻对纵向数据的回归性能
Mohammad Sadegh Loeloe1, Seyyed Mohammad Tabatabaei2,3, Reyhane Sefidkar1
1Center for Healthcare Data Modeling, Department of Biostatistics and Epidemiology, School of Public Health, Shahid Sadoughi University of Medical Sciences, Yazd, Iran.
BMC bioinformatics
|October 1, 2025
概括
基于集群的KNN纵向数据回归 (CKNNRLD) 提高了纵向数据分析的预测准确性和效率. 这种新的方法优于标准的K-近邻 (KNN) 回归,特别是对于大型数据集.
科学领域:
- 统计和数据科学 统计和数据科学
- 生物统计学 生物统计学
- 机器学习 机器学习
背景情况:
- 纵向研究需要灵活的预测方法来预测响应轨迹.
- 时间依赖和时间独立的共变量在纵向数据分析中带来了挑战.
- 现有的K-Nearest Neighbor (KNN) 回归可能会与大型纵向数据集作斗争.
研究的目的:
- 引入基于集群的KNN纵向数据回归 (CKNNRLD),这是KNN的一种新扩展.
- 为了提高纵向数据的预测准确性和计算效率.
- 为分析复杂的纵向数据集提供一个强大的工具.
主要方法:
- 使用K-means用于纵向数据 (KML) 算法进行数据聚类.
- 最接近邻居的搜索仅限于相关数据集群.
- 通过广泛的模拟和真实螺旋测量数据集来开发和验证理论框架.
主要成果:
- 与标准KNN相比,CKNNRLD显示出更高的预测准确度.
- CKNNRLD显著减少了执行时间和计算负担.
- 对于N=2000,CKNNRLD的速度大约是标准KNN的3.7倍,对于N>100和N>500的速度有显著的改善.
结论:
- 与传统的KNN相比,CKNNRLD在准确性和计算效率方面提供了显著的改进.
- 该算法对于管理大型纵向数据集的研究人员来说尤其有益.
- CKNNRLD为纵向数据预测提供了有价值的进步.
相关概念视频
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Residuals and Least-Squares Property
9.1K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.1K
Longitudinal Studies
479
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
479
Longitudinal Research
13.1K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
13.1K
