Related Experiment Video
Updated: Jan 31, 2026

08:50
Predictive Immune Modeling of Solid Tumors
Published on: February 25, 2020
7.6K
Penalized negative binomial models for modeling an overdispersed count outcome with a high-dimensional predictor
Rebecca R Lehman1, Kellie J Archer2
1United Network for Organ Sharing, Richmond, VA, United States of America.
Plos One
|January 9, 2019
Summary
This study introduces a new penalized negative binomial regression model for analyzing micronuclei (MN) frequency data, especially when the number of variables exceeds samples. The model handles over-dispersed count data, improving genotoxicity and cancer risk assessments.
Area of Science:
- Genomics
- Toxicology
- Biostatistics
Background:
- Micronuclei (MN) are biomarkers for genotoxic exposure and cancer risk.
- Existing MN scoring guidelines lack statistical analysis methods, leading to varied interpretations.
- Integrating MN data with high-throughput genomic technologies can reveal molecular features of micronucleation.
Purpose of the Study:
- To address the lack of statistical methods for analyzing MN frequency data, particularly in high-dimensional settings.
- To present a novel penalized negative binomial regression model for over-dispersed count data when the number of explanatory variables (P) exceeds the number of samples (N).
- To evaluate the performance of the proposed model against penalized Poisson models using simulation studies.
Main Methods:
- Developed a penalized negative binomial regression model for count data analysis where P > N.
- Conducted simulation studies to compare the new model with penalized Poisson models for over-dispersed outcomes.
- Applied the model to analyze micronuclei frequency and gene expression data from the Norwegian Mother and Child Cohort Study.
Main Results:
- The penalized negative binomial regression model effectively handles over-dispersed count data in high-dimensional settings (P > N).
- Demonstrated superior performance compared to penalized Poisson models under over-dispersion.
- Successfully applied the method to real-world MN frequency and gene expression data.
Conclusions:
- The developed penalized negative binomial regression model provides a robust statistical framework for analyzing MN frequency data, especially with high-dimensional genomic information.
- The `countgmifs` R package facilitates the application of this method to discrete outcome datasets with high-dimensional covariates.
- This approach enhances the statistical rigor for interpreting genotoxic exposure and cancer risk biomarkers.
Related Concept Videos
Predicting Reaction Outcomes
10.8K
Kinetics describes the rate and path by which a reaction occurs. In contrast, thermodynamics deals with state functions and describes the properties, behavior, and components of a system. It is not concerned with the path taken by the process and cannot address the rate at which a reaction occurs. Although it does provide information about what can happen during a reaction process, it does not describe the detailed steps of what appears on an atomic or a molecular level. On the other hand,...
10.8K
The Binomial Theorem
322
The Binomial Theorem is a foundational principle in algebra used to expand expressions raised to a power. It provides a structured approach for expanding binomials of the form (a+b)n, where a and b are variables or constants representing algebraic expressions, and n is a non-negative integer.The general form of the Binomial Theorem is:Each term in the expansion involves a binomial coefficient, which is calculated using factorials:The exponent of a in each term decreases from n to 0, while the...
322
Molecular Models
43.7K
Physical models representing molecular architectures of chemical compounds play essential roles in understanding chemistry. The use of molecular models makes it easier to visualize the structures and shapes of atoms and molecules.
43.7K
Binomial Probability Distribution
15.9K
A binomial distribution is a probability distribution for a procedure with a fixed number of trials, where each trial can have only two outcomes.
The outcomes of a binomial experiment fit a binomial probability distribution. A statistical experiment can be classified as a binomial experiment if the following conditions are met:
There are a fixed number of trials. Think of trials as repetitions of an experiment. The letter n denotes the number of trials.
There are only two possible outcomes,...
The outcomes of a binomial experiment fit a binomial probability distribution. A statistical experiment can be classified as a binomial experiment if the following conditions are met:
There are a fixed number of trials. Think of trials as repetitions of an experiment. The letter n denotes the number of trials.
There are only two possible outcomes,...
15.9K
Binomial Expansion Using Pascal's Triangle
256
Expanding a binomial expression such as (a + b)n results in a predictable sequence of terms that can be systematically derived using Pascal’s Triangle. This triangular array of numbers plays a central role in understanding and computing the coefficients of binomial expansions.Pascal’s Triangle is constructed such that each row corresponds to the coefficients of a binomial raised to a power. The topmost row, known as the zeroth row, corresponds to (a + b)0, and each successive row...
256
Predicting Molecular Geometry
45.8K
VSEPR Theory for Determination of Electron Pair Geometries
45.8K

