用于复合量子回归的贝叶斯增量树集
Yaeji Lim1, Ruijin Lu2, Madeleine St Ville3
1Department of Applied Statistics, Chung-Ang University, Seoul, Korea.
概括
这项研究引入了复合量子BART,一种用于建模复杂数据关系的新统计方法. 它提供了更好的预测准确性,特别是在不寻常的错误分布下,优于现有的技术.
科学领域:
- 统计数据
- 机器学习
- 经济计量学
背景情况:
- 传统的定量回归模型是特定的定量.
- 贝叶斯增量回归树 (BART) 处理复杂的非线性关系.
- 复合量子回归 (CQR) 提供了错误分布的稳定性.
研究的目的:
- 开发一种结合BART和CQR的新统计方法.
- 在多种错误分布下增强复杂的预测结果关系的建模.
- 与现有方法相比,提高预测性能.
主要方法:
- 贝叶斯增量回归树 (BART) 与复合量子回归 (CQR) 的整合.
- 开发一种灵活的方法来捕捉响应变量的全部条件分布.
- 使用BART的非线性建模和CQR的稳定性.
主要成果:
- 拟议的复合量子BART方法显示出卓越的预测性能.
- 优于经典BART,量子BART和复合量子线性回归模型.
- 实现显著的根平均平方误差 (RMSE) 减少,特别是在重尾或受污染的误差分布下.
结论:
- 复合量子BART为统计建模提供了强大而灵活的方法.
- 该方法对于具有非标准错误分布的数据集特别有利.
- 在预测准确度上提供了实质性的改进.
更多相关视频
相关概念视频
Survival Tree
159
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
159
Parametric Survival Analysis: Weibull and Exponential Methods
600
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
600
Distributions to Estimate Population Parameter
4.3K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.3K
Binomial Probability Distribution
11.4K
A binomial distribution is a probability distribution for a procedure with a fixed number of trials, where each trial can have only two outcomes.
The outcomes of a binomial experiment fit a binomial probability distribution. A statistical experiment can be classified as a binomial experiment if the following conditions are met:
There are a fixed number of trials. Think of trials as repetitions of an experiment. The letter n denotes the number of trials.
There are only two possible outcomes,...
The outcomes of a binomial experiment fit a binomial probability distribution. A statistical experiment can be classified as a binomial experiment if the following conditions are met:
There are a fixed number of trials. Think of trials as repetitions of an experiment. The letter n denotes the number of trials.
There are only two possible outcomes,...
11.4K
Contingency Table
2.6K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
2.6K
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K


