Least squares regression methods for clustered ROC data with discrete covariates.
Liansheng Larry Tang1,2, Wei Zhang3, Qizhai Li3
1Department of Statistics, George Mason University, Fairfax, VA 22030, USA.
Biometrical Journal. Biometrische Zeitschrift
|February 6, 2016
Summary
This study introduces novel least squares methods for Receiver Operating Characteristic (ROC) curve estimation in clustered diagnostic test data. These methods improve efficiency and handle complex data structures for accurate diagnostic test evaluation.
Area of Science:
- Biostatistics
- Medical Diagnostics
- Statistical Modeling
Background:
- Receiver Operating Characteristic (ROC) curves are crucial for evaluating diagnostic test accuracy with continuous or ordinal results.
- Clustered data from multiple tests on the same subject present challenges for traditional ROC analysis.
- Existing least squares methods for ROC curves do not adequately address clustered data or its statistical properties.
Purpose of the Study:
- To develop and evaluate least squares ROC methods specifically designed for clustered data with discrete covariates.
- To investigate the statistical properties and efficiency of these novel methods compared to existing nonparametric approaches.
- To provide a robust tool for accurate diagnostic test evaluation in complex, clustered data settings.
Main Methods:
- Development of least squares ROC methods allowing flexible baseline and link functions.
- Adaptation of methods to accommodate clustered data structures and discrete covariates.
- Derivation of asymptotic properties for the proposed statistical methods.
Main Results:
- The proposed least squares methods generate smooth ROC curves that respect the continuous nature of underlying true curves.
- Simulation studies demonstrate superior efficiency of the least squares methods over nonparametric ROC methods under specific model assumptions.
- The methods were successfully applied to a real-world example of detecting glaucomatous deterioration.
Conclusions:
- The developed least squares ROC methods effectively handle clustered data and discrete covariates, offering an advancement in diagnostic test evaluation.
- These methods provide a more efficient and statistically sound approach compared to existing techniques for complex datasets.
- The findings have significant implications for accurately assessing diagnostic test performance in clinical research and practice.
Related Concept Videos
Residuals and Least-Squares Property
9.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.8K
Multiple Regression
4.3K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.3K
Receiver Operating Characteristic Plot
570
A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
570
Friedman Two-way Analysis of Variance by Ranks
556
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
556
Regression Analysis
8.9K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.9K
Cluster Sampling Method
15.5K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.5K


