用敲除过器增强梯度:一种生物统计方法来对变量选择.
1College of Health Sciences, The University of Memphis, 3720 Alumni Ave, Memphis, TN, 38152, USA. A.Mohamed@memphis.edu.
BMC bioinformatics
|November 26, 2025
概括
本研究引入了一种新的可变选择方法,将仿制物与光梯度增强机 (LightGBM) 和SHAP值集成在一起. 该方法有效地识别了重要的变量,同时控制了错误发现率 (FDR),在大数据场景中表现优于传统方法.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 随着数据复杂性的增加,需要有效的变量选择方法.
- 控制虚假发现率 (FDR) 和保持统计能力是关键的挑战.
- 淘汰过器提供了一个强大的方法,通过创建用于推断的负控制.
研究的目的:
- 扩大对光梯度增强机 (LightGBM) 的仿制过器的应用.
- 在大数据的背景下,提高变量选择准确性和效率.
- 用SHAP值来提高机器学习模型的可解释性.
主要方法:
- 将淘汰变量生成与LightGBM集成.
- 使用形状添加式解释 (SHAP) 进行模型解释性.
- 为了验证,进行了广泛的实验和模拟研究.
主要成果:
- 拟议的方法准确地识别了每个类别的重要变量.
- 与传统方法相比,证明了卓越的性能,速度和效率.
- 通过SHAP值来提高变量重要性的解释性.
结论:
- 将仿制品集成到LightGBM中,为变量选择提供了一个强大的工具.
- 这种方法有效地解决了大数据分析的挑战.
- 该方法通过提高性能和可解释性来推进统计建模和机器学习应用.
相关概念视频
Bootstrapping
792
The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is...
792
Quantifying and Rejecting Outliers: The Grubbs Test
3.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K
Biostatistics: Overview
705
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
705
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
1.1K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
1.1K
Frequency-dependent Selection
23.0K
When the fitness of a trait is influenced by how common it is (i.e., its frequency) relative to different traits within a population, this is referred to as frequency-dependent selection. Frequency-dependent selection may occur between species or within a single species. This type of selection can either be positive—with more common phenotypes having higher fitness—or negative, with rarer phenotypes conferring increased fitness.
23.0K
Prediction Intervals
3.1K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.1K


