Clustered data analysis under miscategorized ordinal outcomes and missing covariates
Surupa Roy1, Subrata Rana2, Kalyan Das2
1Department of Statistics, St. Xavier's College, Kolkata, India.
Statistics in Medicine
|July 29, 2015
Summary
This study introduces a flexible clustered ordinal model to analyze data with missing covariates and miscategorized outcomes. The novel two-step approach effectively handles common data imperfections in statistical modeling.
Area of Science:
- Statistics
- Biostatistics
- Data Analysis
Background:
- Ordinal data analysis is crucial in many scientific fields.
- Real-world data frequently exhibit missing covariate information and outcome miscategorization.
- Existing models may not adequately address these common data complexities.
Purpose of the Study:
- To develop and analyze a clustered ordinal model accommodating missing covariates.
- To address the challenge of miscategorized data using surrogate variables.
- To investigate the influence of age, sex, and food habits on plaque deposit using orthodontic data.
Main Methods:
- A general model structure is proposed to incorporate information from surrogate variables.
- A novel two-step estimation approach is introduced for model parameter estimation.
- The model's flexibility allows it to handle both missingness and miscategorization.
Main Results:
- The proposed model effectively handles clustered ordinal data with missing covariates.
- The two-step estimation method provides a robust way to estimate model parameters.
- Simulation studies validate the model's performance and applicability.
Conclusions:
- The developed clustered ordinal model offers a flexible solution for analyzing complex datasets.
- The proposed methodology is suitable for various applications, including orthodontic research.
- The approach effectively mitigates the impact of missing data and miscategorization on analysis.
More Related Videos
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
711
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
711
Ordinal Level of Measurement
37.5K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
37.5K
Friedman Two-way Analysis of Variance by Ranks
578
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
578
Contingency Table
5.0K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
5.0K
One-Way ANOVA: Unequal Sample Sizes
7.0K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
7.0K
How Data are Classified: Categorical Data
48.4K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
48.4K


