Related Experiment Video
Updated: Aug 15, 2025

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.2K
Diagnosing and Handling Common Violations of Missing at Random
Feng Ji1, Sophia Rabe-Hesketh2, Anders Skrondal3,4,5
1University of California, Berkeley University of Toronto, Berkeley, USA.
Psychometrika
|January 4, 2023
Summary
This study introduces diagnostic tests to address violations of the missing at random (MAR) assumption in statistical modeling. A novel test-based estimator improves handling of severe missing data problems by selecting appropriate data-deletion methods.
Area of Science:
- Statistics
- Econometrics
- Machine Learning
Background:
- Ignorable likelihood (IL) methods are standard for handling missing data in multivariate models under the missing at random (MAR) assumption.
- Violations of MAR, where unobserved variables influence missingness, necessitate advanced techniques.
- Existing data-deletion methods, while effective, require careful selection based on the MAR violation type.
Purpose of the Study:
- To develop diagnostic tests for identifying specific violations of the MAR assumption.
- To propose a test-based estimator that adaptively applies data-deletion strategies.
- To evaluate the performance of the proposed estimator against standard IL approaches.
Main Methods:
- Development of a likelihood-ratio test for heteroscedastic regression models.
- Implementation of a kernel conditional independence test for MAR violation detection.
- Construction of a test-based estimator utilizing these diagnostic tools.
Main Results:
- The proposed diagnostic tests effectively identify MAR violations.
- The test-based estimator demonstrates superior performance over IL methods in severe missing data scenarios.
- The estimator shows comparable performance to IL methods when missing data issues are less pronounced.
Conclusions:
- The developed diagnostic tests and test-based estimator offer a robust solution for handling MAR violations in statistical modeling.
- This approach provides a data-driven method for selecting appropriate missing data handling strategies.
- The findings have significant implications for improving the accuracy and reliability of statistical analyses with incomplete datasets.
Related Concept Videos
Random and Systematic Errors
11.6K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
11.6K
Random Error
1.3K
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
1.3K
Types of Errors: Detection and Minimization
1.9K
Error is the deviation of the obtained result from the true, expected value or the estimated central value. Errors are expressed in absolute or relative terms.
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
1.9K
Unusual Results
3.3K
Unusual results are those that have a very low chance of occurring. Unusual results can be identified using probabilities and the range rule of thumb. In problems involving probability, unusual results can be observed in 2 instances – an unusually high number of successes or an unusually low number of successes.
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
3.3K
Wald-Wolfowitz Runs Test II
297
The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
297
Wald-Wolfowitz Runs Test I
703
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
The test works...
703

