Stata Modules for Calculating Novel Predictive Performance Indices for Logistic Models.
Mahnaz Barkhordari1, Mojgan Padyab2, Farzad Hadaegh3
1Department of Mathematics, Bandar Abbas Branch, Islamic Azad University, Bandar Abbas, IR Iran.
International Journal of Endocrinology and Metabolism
|June 10, 2016
Summary
A new Stata command, addpred, simplifies assessing how new biomarkers improve cardiovascular disease (CVD) risk prediction. This tool aids researchers in evaluating predictive model performance using methods like Net Reclassification Improvement (NRI) and Integrated Discriminatory Improvement (IDI).
Area of Science:
- Biostatistics
- Cardiovascular Disease Epidemiology
- Medical Informatics
Background:
- Cardiovascular disease (CVD) risk prediction is crucial for prevention, with numerous biomarkers emerging.
- Assessing the added predictive value of novel biomarkers requires robust statistical methods like discrimination, calibration, and reclassification indices.
- Limited availability of user-friendly software hinders the widespread adoption of these advanced model assessment techniques.
Purpose of the Study:
- To develop a user-friendly statistical software tool for researchers with limited programming expertise.
- To facilitate the implementation of novel methods for assessing the performance of risk prediction models.
- To enable quantification of the improvement in risk prediction offered by new biomarkers.
Main Methods:
- Developed a Stata command named 'addpred' for logistic regression models.
- The command calculates cut-point-free and cut-point-based Net Reclassification Improvement (NRI) and Integrated Discriminatory Improvement (IDI).
- Applied the command to real-world data from the Tehran Lipid and Glucose Study (TLGS) to evaluate the Framingham CVD risk algorithm.
Main Results:
- The 'addpred' Stata command is available for logistic regression models.
- Demonstrated the utility of the command in assessing the predictive performance improvement using family history, waist circumference, and glucose levels.
Conclusions:
- The 'addpred' Stata package simplifies the assessment of novel biomarker predictive capacity.
- This tool is expected to encourage the broader application of advanced statistical methods in biomarker research.
- Facilitates more accurate cardiovascular risk prediction and prevention strategies.
Related Concept Videos
Regression Analysis
8.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.8K
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
7.0K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
7.0K
Statistical Analysis: Overview
16.9K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
16.9K
Statistical Analysis System (SAS)
1.2K
SAS, short for Statistical Analysis System, is a powerful data analysis, management, and visualization tool. Developed by the SAS Institute in the early 1970s, SAS has evolved into a comprehensive software suite used across various industries for statistical analysis, business intelligence, and predictive modeling.
Applications: SAS finds applications in numerous fields, including healthcare for clinical trial analysis, finance for risk assessment, marketing for customer data analysis, and...
Applications: SAS finds applications in numerous fields, including healthcare for clinical trial analysis, finance for risk assessment, marketing for customer data analysis, and...
1.2K
Goodness-of-Fit Test
9.4K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
9.4K
Prediction Intervals
3.5K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.5K


