Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Odds Ratio01:09

Odds Ratio

396
The odds ratio (OR) is a statistical measure used extensively in epidemiology and research to quantify the strength of association between exposure and outcome across different groups. Unlike relative risk, which compares the probabilities of an event occurring, the odds ratio compares the odds of an event occurring in the exposed group to the odds of it occurring in the unexposed group. The odds, in this context, are calculated as the probability of the event happening divided by the...
396
Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

809
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:  
809
Bias01:22

Bias

6.4K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
6.4K
Assumptions of Survival Analysis01:15

Assumptions of Survival Analysis

218
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
218
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

8.1K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.1K
Survival Tree01:19

Survival Tree

183
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
183

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Hypertension and Postoperative Outcomes: A Retrospective Cohort Study.

Anesthesiology research and practice·2026
Same author

Kidney and Cardiovascular Outcomes by CKD stage and etiology of CKD: A Randomly Selected Nationwide Prospective Multicenter Cohort Study in Japan (Reach-J Study).

American journal of nephrology·2026
Same author

Association of the PD-L1 CPS With the Efficacy of First-Line Nivolumab Plus Chemotherapy for Unresectable Advanced or Recurrent Gastric Cancer.

Cancer medicine·2026
Same author

Measurement properties of the interest in health scale among community-dwelling older adults in Japan: Verification of the 12-item, 6-item, and 4-item versions of the interest in health scale.

Preventive medicine reports·2026
Same author

Balloon-occluded Alternative Infusion of Fragmented Gelatin Particles of TACE for Hepatocellular Carcinoma Refractory to Atezolizumab-Bevacizumab.

Anticancer research·2026
Same author

Mid-treatment MRI-based tumor volume reduction rate as a continuous prognostic factor after chemoradiation for cervical cancer: development and two-center internal-external validation.

Journal of radiation research·2026

Related Experiment Video

Updated: Oct 19, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.7K

Bias in Odds Ratios From Logistic Regression Methods With Sparse Data Sets.

Masahiko Gosho1, Tomohiro Ohigashi2,3, Kengo Nagashima4

  • 1Department of Biostatistics, Faculty of Medicine, University of Tsukuba.

Journal of Epidemiology
|September 27, 2021
PubMed
Summary

Sparse data bias in logistic regression can lead to inaccurate odds ratios (ORs). Bayesian methods, particularly those using log F-type or hyper-ɡ priors, offer superior bias reduction compared to classical methods for sparse data.

Keywords:
Bayesian methodsFirth’s penalizationexact logistic regression methodɡ-prior

More Related Videos

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.2K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

269

Related Experiment Videos

Last Updated: Oct 19, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.7K
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.2K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

269

Area of Science:

  • Statistics
  • Epidemiology
  • Biostatistics

Background:

  • Logistic regression models are crucial for analyzing binary outcomes but suffer from sparse data bias.
  • This bias, stemming from limited participants, inflates odds ratios (ORs) and is often overlooked in epidemiological studies.
  • Sparse data bias can lead to unreliable and excessively large OR estimates.

Approach:

  • This study evaluates Bayesian methods against classical approaches (Maximum Likelihood, Firth's, exact methods) for reducing sparse data bias.
  • A simulation study compares the performance of these methods.
  • The methods were also applied to a real-world dataset.

Key Points:

  • Simulation results show considerable bias in Maximum Likelihood, Firth's, and exact methods.
  • Bayesian methods with hyper-ɡ priors reduced bias under the null hypothesis.
  • Bayesian methods with log F-type priors effectively reduced bias under the alternative hypothesis.

Conclusions:

  • Bayesian methods with log F-type and hyper-ɡ priors outperform classical methods for sparse logistic regression.
  • The optimal method choice depends on whether the null or alternative hypothesis is of primary interest.
  • Sensitivity analysis is crucial for validating results in sparse data scenarios.