Local linear smoothing for regression surfaces on the simplex using Dirichlet kernels
Christian Genest1, Frédéric Ouimet1
1Department of Mathematics and Statistics, McGill University, 805, rue Sherbrooke ouest, Montréal, QC H3A 0B9 Canada.
Summary
This study presents a new local linear smoother for simplex regression surfaces. The novel Dirichlet kernel estimator demonstrates superior performance compared to existing methods in simulation studies.
Area of Science:
- Statistics
- Nonparametric Regression
- Computational Statistics
Background:
- Regression analysis on the simplex is crucial for modeling compositional data.
- Existing methods often struggle with boundary properties.
- Local polynomial smoothing offers improved performance over local constant methods.
Purpose of the Study:
- To introduce a novel local linear smoother for regression on the simplex.
- To analyze the asymptotic properties of the proposed estimator.
- To compare its performance against existing estimators.
Main Methods:
- Developing a local linear smoother using a weighted least-squares approach.
- Employing a locally adaptive Dirichlet kernel for weighting.
- Deriving asymptotic results for bias, variance, and mean squared error.
- Conducting simulation studies for performance evaluation.
Main Results:
- The proposed local linear smoother exhibits favorable boundary properties.
- Asymptotic properties (bias, variance, MSE, MISE) are theoretically established.
- Simulation results indicate the new estimator outperforms the Nadaraya-Watson estimator with a Dirichlet kernel.
Conclusions:
- The local linear smoother with a Dirichlet kernel is an effective method for regression on the simplex.
- The theoretical and simulation results support its practical applicability.
- This work extends univariate smoothing results to the multivariate simplex domain.
Keywords:
Adaptive estimatorAsymmetric kernelBeta kernelBoundary biasDirichlet kernelLocal linear smootherMean integrated squared errorNadaraya–Watson estimatorNonparametric regressionRegression surfaceSimplexMore Related Videos
Related Concept Videos
Linear Approximation in Frequency Domain
80
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
80
Linear Approximation in Time Domain
59
Nonlinear systems often require sophisticated approaches for accurate modeling and analysis, with state-space representation being particularly effective. This method is especially useful for systems where variables and parameters vary with time or operating conditions, such as in a simple pendulum or a translational mechanical system with nonlinear springs.
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
59
Curvilinear Motion: Rectangular Components
382
Curvilinear motion characterizes the movement of a particle or object along a curved path, notably evident when envisioning a car navigating a winding road. If the car starts at point A, its position vector is established within a fixed frame of reference, where the ratio of the position vector to its magnitude signifies the unit vector pointing in the position vector's direction.
As the car advances, its position evolves over time. Quantifying the car's velocity involves computing the...
As the car advances, its position evolves over time. Quantifying the car's velocity involves computing the...
382
Residuals and Least-Squares Property
7.2K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.2K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
34
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
34
Calibration Curves: Linear Least Squares
1.2K
A calibration curve is a plot of the instrument's response against a series of known concentrations of a substance. This curve is used to set the instrument response levels, using the substance and its concentrations as standards. Alternatively, or additionally, an equation is fitted to the calibration curve plot and subsequently used to calculate the unknown concentrations of other samples reliably.
For data that follow a straight line, the standard method for fitting is the linear...
For data that follow a straight line, the standard method for fitting is the linear...
1.2K


