Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Sign Test for Matched Pairs01:17

Sign Test for Matched Pairs

The sign test for matched pairs offers a robust method for comparing two paired samples, often for the effects of an intervention in one of them. This method is very useful in situations where the underlying distribution of the data is unknown. The test compares two related samples—often pre- and post-treatment measurements on the same subjects—to determine if there are significant differences in their median values.
To conduct the sign test, we first calculate the differences in value between...
Regression Toward the Mean01:52

Regression Toward the Mean

Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when researchers try to extrapolate results...
Cluster Sampling Method01:20

Cluster Sampling Method

Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Statistical Hypothesis Testing01:16

Statistical Hypothesis Testing

Hypothesis testing is a critical statistical procedure facilitating informed, evidence-based decisions. It begins with a hypothesis, which is a tentative explanation, or a prediction about a population parameter. This hypothesis can be either a null hypothesis (H0), indicating no effect or difference, or an alternative hypothesis (Ha), suggesting an effect or difference.
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test01:09

Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test

In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with data...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Perfluorooctanoic acid and cancer incidence: an updated investigation of a cohort in the mid-Ohio Valley.

Environment international·2026
Same author

Circulating pre-diagnostic metabolites and risk of hepatocellular carcinoma and intrahepatic cholangiocarcinoma: a population-based study of 12 cohorts.

Journal of the National Cancer Institute·2026
Same author

Artificially Sweetened and Sugar-Sweetened Beverage Intake and Risk of Liver Cancer.

JAMA network open·2026
Same author

Cessation of Betel Quid Chewing, Smoking, and Alcohol Drinking and Risk of Oral Precancer and Oral Cancer.

JCO global oncology·2026
Same author

Sex differences in cancer incidence persist across race and ethnicity.

Biology of sex differences·2026
Same author

Sociodemographic characteristics of populations living near industrial land disposals of known and suspected carcinogens across the United States.

Journal of exposure science & environmental epidemiology·2026

Related Experiment Videos

Testing logistic regression coefficients with clustered data and few positive outcomes.

Sally Hunsberger1, Barry I Graubard, Edward L Korn

  • 1Biometric Research Branch, National Cancer Institute, Bethesda, MD 20892, U.S.A. sallyh@ctep.nci.nih.gov

Statistics in Medicine
|August 21, 2007
PubMed
Summary

A new simulation-based method improves logistic regression analysis for clustered data with few positive outcomes. This approach offers more accurate hypothesis testing than existing methods, especially in health research.

Related Experiment Videos

Area of Science:

  • Biostatistics
  • Epidemiology
  • Statistical modeling

Background:

  • Logistic regression is common for analyzing clustered data, but struggles with sparse positive outcomes in certain categories.
  • Existing methods like generalized Wald, score tests, and bootstrap tests can yield unreliable results (inflated or conservative levels) with few positive outcomes.
  • Accurate statistical analysis is crucial for identifying associations, such as asthma risk factors in large health surveys.

Purpose of the Study:

  • To develop and evaluate a robust simulation-based method for hypothesis testing in logistic regression with clustered samples, specifically addressing the challenge of few positive outcomes.
  • To compare the performance of the proposed method against established tests (generalized Wald, score, bootstrap) in maintaining nominal significance levels.

Main Methods:

  • The study proposes a novel simulation-based methodology for testing logistic regression coefficients in the presence of sparse data within cluster samples.
  • Performance was evaluated through simulations, assessing the method's ability to maintain nominal levels compared to existing techniques.
  • The proposed method's utility was also examined for logistic regression model goodness-of-fit testing using deciles-of-risk tables.

Main Results:

  • Simulations demonstrated that traditional tests (generalized Wald, score, bootstrap) can exhibit unstable performance (inflated or conservative levels) when positive outcomes are few.
  • The proposed simulation-based method effectively maintains nominal levels, outperforming existing methods in scenarios with sparse positive outcomes.
  • The new methodology also proved beneficial for assessing the goodness-of-fit of logistic regression models.

Conclusions:

  • A new simulation-based testing methodology provides a more reliable approach for logistic regression analysis with clustered data, particularly when dealing with few positive outcomes.
  • This method offers improved accuracy and stability compared to generalized Wald, score, and bootstrap tests.
  • The proposed approach enhances the validity of statistical inference in health research and model diagnostics.