混合效应调频边界普通森林:用层次数据进行普通预测的树木组合方法
1Department of Statistics, TU Dortmund University, Dortmund, Germany.
Multivariate behavioral research
|September 9, 2025
概括
一种新的机器学习方法,混合效应调整频率边界顺序森林 (mixfabOF),改善了对等级数据的顺序预测. 它的性能优于现有方法,特别是当随机效应具有高可变性时.
科学领域:
- 社会和生命科学 社会和生命科学
- 统计建模 统计建模
- 机器学习 机器学习
背景情况:
- 在社会和生命科学中,顺序预测对于分析学校成绩或评分表等任务至关重要.
- 现有的机器学习 (ML) 方法,如随机森林 (RF),显示出高的预测能力,但往往缺乏对层次数据结构的支持.
- 在这些领域常见的等级数据 (例如,课堂上的学生) 需要专门的方法来准确预测.
研究的目的:
- 扩展频率调整边界普通森林 (fabOF) 机器学习方法以适应层次数据.
- 引入混合效应调节频率的边界顺序森林 (mixfabOF),以改善嵌套数据设置中的顺序预测.
- 为了评估混合fabOF与现有的基于RF的顺序预测方法的性能.
主要方法:
- 通过扩展 fabOF 方法,开发混合效应调频边界普通森林 (mixfabOF).
- 在混合fabOF框架内使用代期望-最大化类型的估计程序.
- 在层次数据集上对混合fabOF与fabOF和其他基于RF的顺序预测技术进行比较分析.
主要成果:
- 混合fabOF在具有高随机效应可变性的层次数据设置中,与fabOF和其他基于射频的方法相比,表现优越.
- 对于具有较低随机效应变量的设置,mixfabOF的性能与fabOF和基于RF的替代方法相美.
- 拟议的方法有效地处理在顺序预测任务中嵌套数据结构的复杂性.
结论:
- 混合效应频率调整边界顺序森林 (mixfabOF) 为层次数据中的顺序预测提供了重大进展.
- 该方法为处理嵌套数据结构的社会和生命科学研究人员提供了有价值的工具.
- mixfabOF提高了预测准确度,特别是在具有大量随机效应变化的场景中.
更多相关视频
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
468
12:26Integrating Remote Sensing with Species Distribution Models; Mapping Tamarisk Invasions Using the Software for Assisted Habitat Modeling SAHM
Published on: October 11, 2016
13.8K
相关概念视频
Ordinal Level of Measurement
31.8K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
31.8K
Survival Tree
362
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
362
Expected Frequencies in Goodness-of-Fit Tests
7.0K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
7.0K
Friedman Two-way Analysis of Variance by Ranks
465
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
465
Prediction Intervals
3.1K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.1K
Ranks
436
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
436
