新的二次差异分析算法用于相关的听力测量数据.
Fuyu Guo1, David M Zucker2, Kenneth I Vaden3
1Department of Epidemiology, Harvard T.H. School of Public Health, Boston, Massachusetts, USA.
Statistics in medicine
|October 26, 2024
概括
新的算法利用配对器官的相关性,比如人类的耳朵,以改善疾病预测模型. 与传统方法相比,这种方法可以提高听力测量表型的准确性.
科学领域:
- 生物统计学 生物统计学
- 医疗信息学 医疗信息学
- 遗传学 是一个遗传学.
背景情况:
- 配对的器官 (眼睛,耳朵,肺) 显示相关的数据.
- 现有的模型经常忽略这些相关性,可能会丢失信息.
- 这特别适用于听力测量现象型预测.
研究的目的:
- 开发新的二次差异分析 (QDA) 算法,以计算对象器官之间的数据依赖.
- 通过利用器官之间的相关性来提高分类模型在预测疾病表型方面的性能.
- 解决传统方法的局限性,即将配对器官数据视为独立的.
主要方法:
- 提出了两阶段的分析策略:数据转换和新的QDA算法.
- 开发了QDA算法,以部分利用两个耳朵的表型之间的依赖性.
- 进行模拟研究,并将算法应用于队列研究中的听力测量数据.
主要成果:
- 数据转换的好处在较小的样本大小中最为明显.
- 拟议的QDA算法比传统方法表现出更高的性能.
- 在人级和耳级都实现了更高的准确性.
结论:
- 对配对器官数据中的相关性进行核算可以显著提高预测模型的准确性.
- 开发的PairQDA R包为这些先进的算法提供了实际实现.
- 这项工作为生物医学研究中分析配对器官数据提供了更有效的方法.
相关概念视频
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Calculating and Interpreting the Linear Correlation Coefficient
5.9K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
5.9K
Detection of Gross Error: The Q Test
5.6K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
5.6K
Linear Approximation in Frequency Domain
86
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
86
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K


