提高基于回归的高维数据集集群的准确性和内部一致性
Bo Zhang1, Jianghua He1, Jinxiang Hu1
1Department of Biostatistics & Data Science, University of Kansas Medical Center, Kansas City, KS 66160, USA.
概括
这项研究通过取代其特征选择方法,增强了对高维数据的组件智能稀少混合回归 (CSMR). 适应式-拉索集成显著提高了集群精度和内部一致性.
科学领域:
- 生物信息学是一种生物信息学.
- 统计遗传学 统计遗传学
- 计算生物学 计算生物学
背景情况:
- 分子智能稀混合回归 (CSMR) 是一种有前途的方法,用于识别分子数据和连续表型之间的复杂关系.
- 由于其特征选择技术的局限性,现有的CSMR方法可能会产生与高维分子数据不一致的结果.
研究的目的:
- 调查是否用替代的调整回归技术取代CSMR中的默认特征选择方法,可以提高集群精度和内部一致性.
- 用模拟和真实生物数据在高维设置中评估修改后的CSMR算法的性能.
主要方法:
- 这项研究修改了CSMR框架,结合了各种规则化的回归方法:拉索,弹性网,SCAD,MCP和自适应拉索.
- 通过计算真正阳性率 (TPR),真负率 (TNR),内部一致性 (IC) 和聚类精度来评估性能.
- 与原来的CSMR算法进行了比较,使用了广泛的模拟研究和真实生物数据集.
主要成果:
- 与原始方法相比,在CSMR框架内替换Adaptive-Lasso可显著提高内部一致性 (IC) 和集群精度.
- 使用Adaptive-Lasso修改的CSMR表现出强大的性能,即使应用于高维分子数据集.
- 其他经过测试的调节回归方法也显示出潜力,但适应拉索产生了最实质性的改进.
结论:
- 对CSMR方法的修改,特别是Adaptive-Lasso的集成,导致集群性能提高.
- 这些改进的CSMR变体为生物和遗传研究中分析和聚类高维数据集提供了可行的替代方案.
相关概念视频
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Improving Translational Accuracy
11.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.6K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Kendall's Coefficient of Concordance
417
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
417
Accuracy, limits, and approximation
476
Accuracy, limits, and approximations are common in many fields, especially in engineering calculations. These concepts are imperative for ensuring that a given value is as close as possible to its true value.
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
476


