Related Experiment Video
Updated: Mar 27, 2026

04:57
Establishing a Competing Risk Regression Nomogram Model for Survival Data
Published on: October 23, 2020
11.0K
Estimating Success in Predicting a Variable with Nominal Measurement Using Other Variables with Nominal Measurements
Multivariate Behavioral Research
|January 12, 2016
Summary
This study introduces a new method for predicting outcomes using nominal predictor variables. The proposed estimator, pc, combines two biased estimators to provide a more accurate prediction probability (p).
Area of Science:
- Statistics
- Machine Learning
- Predictive Modeling
Background:
- Estimating prediction accuracy is crucial in statistical modeling.
- Existing methods for nominal variables can be biased.
Purpose of the Study:
- To develop and evaluate estimators for prediction success probability (p) with nominal variables.
- To compare the bias and Mean Squared Error (MSE) of different estimators.
Main Methods:
- Proposed three estimators: pa (in-sample success rate), pb (holdout group), and pc (average of pa and pb).
- Conducted simulation studies to assess estimator performance.
- Analyzed bias and MSE of pa, pb, and pc.
Main Results:
- Estimator pa is consistently biased upwards.
- Estimator pb is consistently biased downwards.
- Estimator pc demonstrated improved performance in simulations.
Conclusions:
- The proposed combined estimator pc offers a more reliable estimation of prediction probability.
- pc mitigates the upward bias of pa and downward bias of pb.
- Simulation results support pc as a superior estimator for this prediction task.
Related Concept Videos
Nominal Level of Measurement
41.9K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. Not every statistical operation can be used with every set of data. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal...
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal...
41.9K
Multiple Regression
4.3K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.3K
Variation
8.3K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
8.3K
Sign Test for Nominal Data
446
The sign test is a nonparametric method used to evaluate hypotheses about the median of a single sample or to compare the medians of two related samples. The sign test is particularly useful when dealing with nominal data, which includes distinct categories without an inherent order, such as names, labels, and preferences. Nominal data restricts statistical analysis to evaluating population proportions rather than mean or median values that require continuous data.
For example, consider a...
For example, consider a...
446
Regression Analysis
8.9K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.9K
Survival Tree
468
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
468

