Analysis of Partially Incomplete Tables of Breast Cancer Characteristics with an Ordinal Variable
B Nebiyou Bekele1, Luis E Nieto-Barajas2, Mark F Munsell3
1Gilead Sciences, Inc., USA.
Summary
This study models breast cancer patient characteristics using a dynamic Dirichlet prior model. It estimates joint distributions across disease stages, improving accuracy with aggregated data through meta-analysis.
Area of Science:
- Biostatistics
- Medical Informatics
- Oncology
Background:
- Accurate modeling of clinical characteristics in breast cancer is crucial for patient stratification and treatment.
- Existing methods may struggle with partially aggregated data and dynamic disease progression.
Purpose of the Study:
- To develop a novel statistical model for the joint distribution of four key breast cancer characteristics.
- To dynamically model conditional probabilities of characteristics evolving with disease stage.
- To address challenges posed by partially collapsed data through meta-analysis.
Main Methods:
- Utilized a series of 4x2x2x2 contingency tables with partially collapsed data.
- Proposed a dynamic model using Dirichlet distributions with a Markov prior structure (dynamic Dirichlet prior).
- Employed a data augmentation technique for meta-analysis of aggregated datasets.
Main Results:
- Successfully modeled the joint distribution of estrogen receptor status, nodal involvement, HER2-neu expression, and disease stage.
- The dynamic model effectively captured evolving conditional probabilities across disease stages.
- The data augmentation technique facilitated a robust meta-analysis.
Conclusions:
- The dynamic Dirichlet prior model provides an effective framework for analyzing complex, multi-dimensional clinical data in breast cancer.
- This approach enhances understanding of patient characteristics by "borrowing strength" across disease stages.
- The proposed methods are valuable for meta-analysis of aggregated epidemiological data.
Related Concept Videos
Cancer Survival Analysis
856
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
856
Comparing the Survival Analysis of Two or More Groups
728
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
728
The Mantel-Cox Log-Rank Test
1.3K
The Mantel-Cox log-rank test is a widely used statistical method for comparing the survival distributions of two groups. It tests whether a statistically significant difference exists in survival times between the groups without assuming a specific distribution for the survival data, making it a non-parametric test. This flexibility makes the log-rank test particularly valuable in medical research and other fields where the timing of an event, such as death or disease recurrence, is of...
1.3K
Ordinal Level of Measurement
37.9K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
37.9K
Contingency Table
5.1K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
5.1K
Ranks
584
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
584


