来自具有异质数据的分辨率智能回归的非参数预测分布
Jialu Li1, Wan Zhang2, Peiyao Wang2
1School of Mathematics and Statistics, Beijing Institute of Technology, Beijing 100081, China.
概括
本研究为异质数据引入了一种新的非参数回归方法,提供响应分布而不是单个值. 该方法有效地处理复杂的数据模式,并提供一致的性能,即使在增加尺寸.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 不同质的数据建模对于个性化营销至关重要.
- 现有的回归方法往往侧重于有条件的平均值,可能需要集群信息.
- 解决数据异质性需要先进的统计方法.
研究的目的:
- 提出一种新的非参数分辨率智能回归程序.
- 要估计响应的全部分布,而不仅仅是单个值.
- 为了适应数据异质性而不需要先前的集群信息.
主要方法:
- 使用边际二进制扩展将响应和预测信息分解为分辨率和模式.
- 通过处罚后勤回归来建模分辨率和模式之间的关系.
- 构建一个条件响应直方图来近似分布.
主要成果:
- 拟议的方法提供了响应的估计分布.
- 证明了一个确定的独立性选属性.
- 展现出不断增长的尺寸的一致性.
- 通过模拟和房地产数据集验证的有效性.
结论:
- 解析智能回归为建模异质数据提供了一个强大的工具.
- 与传统方法相比,它提供了对响应变异性的更全面的理解.
- 该方法对于高维度应用来说强大且可扩展.
相关概念视频
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Residual Plots
4.6K
A residual plot is a statistical representation of data used to analyze correlation and regression results. It helps verify the requirements for drawing specific conclusions about correlation and regression. To obtain the residual plot, first, the residual for each data value is calculated, which is simply the vertical distance between the observed and the predicted value obtained from the regression equation.
When the residual values are plotted against the variable x, it is called a residual...
When the residual values are plotted against the variable x, it is called a residual...
4.6K
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
72
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
72
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
133
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
133
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K


