Related Experiment Video
Updated: Jan 3, 2026

04:57
Establishing a Competing Risk Regression Nomogram Model for Survival Data
Published on: October 23, 2020
10.7K
High-dimensional regression with ordered multiple categorical predictors
Lei Huang1, Weiqiang Hang2, Yue Chao1
1School of Mathematics, Southwest Jiaotong University, Chengdu, China.
Statistics in Medicine
|November 29, 2019
Summary
This study introduces a new method for linear regression with ordered multiple categorical (OMC) predictors in high-dimensional data. The approach effectively selects relevant variables and demonstrates strong performance in simulations and real-world analyses.
Area of Science:
- Statistics
- Machine Learning
- Data Science
Background:
- Ordered multiple categorical (OMC) response models are established, but OMC predictors in high-dimensional linear regression are under-researched.
- High-dimensional settings require methods to select relevant variables and handle pseudocategories of discrete predictors.
Purpose of the Study:
- To propose a novel method for linear regression with high-dimensional ordered multiple categorical predictors.
- To address the challenge of automatic variable selection and handling of pseudocategories in OMC predictors.
Main Methods:
- A dummy variable transformation method for OMC predictors.
- An L1 penalty regression approach applied to the transformed variables.
- Derivation of model selection consistency under high-dimensional assumptions.
Main Results:
- The proposed method demonstrates good performance in simulation studies.
- The method shows effectiveness in real data analysis.
- Successful automatic selection of pseudocategories and irrelevant explanatory variables.
Conclusions:
- The developed method offers a robust solution for high-dimensional linear regression with OMC predictors.
- The approach exhibits wide applicability in various regression analysis scenarios.
- Validates the effectiveness of dummy variable transformation and L1 penalty for this specific problem.
Related Concept Videos
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
How Data are Classified: Categorical Data
42.3K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
42.3K
Friedman Two-way Analysis of Variance by Ranks
459
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
459
Multiple Allele Traits
37.8K
The Concept of Multiple Allelism
37.8K
Ordinal Level of Measurement
31.7K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
31.7K
Survival Tree
354
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
354

