Variable selection in finite mixture of regression models using the skew-normal distribution.
Junhui Yin1, Liucang Wu1, Lin Dai1
1Faculty of Science, Kunming University of Science and Technology, Kunming, People's Republic of China.
Journal of Applied Statistics
|June 16, 2022
Summary
This study introduces a new variable selection method for finite mixture of regression (FMR) models using the skew-normal distribution. This approach improves statistical modeling for asymmetric data, offering enhanced accuracy and theoretical properties.
Area of Science:
- Statistics
- Statistical Modeling
- Regression Analysis
Background:
- Variable selection is crucial in finite mixture of regression (FMR) models.
- Standard FMR models often assume normal error distributions, which are inadequate for asymmetric data.
- Existing methods lack robustness when dealing with heterogeneous data exhibiting skewness.
Purpose of the Study:
- To develop a robust variable selection procedure for FMR models accommodating asymmetric error distributions.
- To introduce the use of the skew-normal distribution within the FMR framework for enhanced modeling flexibility.
- To establish the theoretical consistency and oracle properties of the proposed variable selection method.
Main Methods:
- Implementation of a variable selection procedure for FMR models utilizing the skew-normal distribution.
- Development of a modified Expectation-Maximization (EM) algorithm for parameter estimation in the skew-normal FMR models.
- Theoretical analysis to establish consistency in variable selection and the oracle property in estimation.
Main Results:
- The proposed method demonstrates consistency in variable selection, accurately identifying relevant predictors.
- The procedure achieves the oracle property in estimation, indicating efficient parameter estimation.
- Numerical experiments and a real data example validate the effectiveness of the skew-normal FMR approach.
Conclusions:
- The skew-normal distribution provides a flexible and effective alternative for variable selection in FMR models with asymmetric data.
- The developed methodology offers improved statistical modeling capabilities for complex datasets.
- The theoretical properties and empirical validation support the practical utility of this advanced statistical technique.
Related Concept Videos
Choosing Between z and t Distribution
2.9K
The z and the Student t distribution estimate the population mean using the sample mean and standard deviation. However, to decide which distribution to use for a calculation, one needs to determine the sample size, the nature of the distribution, and whether the population standard deviation is known. If the population standard deviation is known and the population is normally distributed, or if the sample size is greater than 30, the z distribution is preferred. The Student t distribution is...
2.9K
Distributions to Estimate Population Parameter
4.3K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.3K
Pharmacokinetic Models: Comparison and Selection Criterion
147
Physiological and compartmental models are valuable tools used in studying biological systems. These models rely on differential equations to maintain mass balance within the system, ensuring an accurate representation of the dynamic processes at play.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
147
Skewness
12.6K
The measures of central tendency calculated from a data set may not reveal much about its intrinsic distribution. If a plot is made of the data set’s values, the mean and the median may not only differ, but also the plot may have more values on one side of the central tendencies. Such a data set is said to be skewed towards that side.
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
12.6K
Types of Skewness
12.5K
If the frequency distribution of a data set is more inclined towards smaller or larger values, the distribution is said to be skewed. If data values are skewed to the right, then the distribution is called positively skewed. Conversely, if the plot is skewed to the left, the distribution is called negatively skewed.
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
12.5K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
100
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
100


