通过残余分析确定偏远群体,并将其应用于医疗保健支出
Hyukdong Kwon1, Jihnhee Yu2, Mingliang Li1
1Department of Economics, University at Buffalo, Buffalo, NY, USA.
Journal of applied statistics
|December 4, 2025
概括
这项研究引入了一种新的数据驱动方法,通过分析购买行为的不明原因变化来识别医疗保健差异群体. 该方法利用外部学生化的余量来确定边缘的消费群体.
科学领域:
- 医疗分析 医疗分析
- 统计建模 统计建模
- 医疗服务研究 医疗服务研究
背景情况:
- 传统的回归分析侧重于整体变量关系,经常忽视无法解释的变量.
- 无法解释的差异可能意味着重要的子组行为或差异.
- 识别这些偏远的行为对于理解市场动态至关重要.
研究的目的:
- 引入一种数据驱动的方法来识别基于偏远行为的差异群体.
- 将这种方法应用于分析医疗保健购买行为,并揭示差异群体的特征.
- 为了利用回归模型中的不明原因变异,获得有针对性的洞察力.
主要方法:
- 开发一种以数据为导向的方法,使用小组学生化余量.
- 计算外部学生化余数的平方平均值,以量化外围行为.
- 对医疗保健市场数据的应用,以检查购买模式.
主要成果:
- 在医疗保健市场中成功确定了特定的差异群体.
- 描述了这些确定的异常群体的独特购买行为.
- 证明了分析无法解释的差异用于差异检测的实用性.
结论:
- 提出的方法有效地通过分析边缘行为来识别差异群体.
- 了解这些群体提供了对医疗保健市场差异的关键见解.
- 这种方法通过关注无法解释的差异来增强传统的回归.
相关概念视频
Outliers and Influential Points
5.9K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
5.9K
What Are Outliers?
4.9K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
4.9K
Detection of Gross Error: The Q Test
6.8K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.8K
Quantifying and Rejecting Outliers: The Grubbs Test
3.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K
Comparing the Survival Analysis of Two or More Groups
538
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
538
Residuals and Least-Squares Property
8.9K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.9K


