Related Experiment Video
Updated: Feb 9, 2026

Mass Spectrometry-Guided Genome Mining as a Tool to Uncover Novel Natural Products
Published on: March 12, 2020
Statistical Power for Postlicensure Medical Product Safety Data Mining
Judith C Maro1,2, Michael D Nguyen3, Inna Dashevsky1,2
1Harvard Medical School.
Sample size calculations for tree-based scan statistics in longitudinal databases ensure robust pharmacovigilance. These methods effectively monitor thousands of health outcomes while controlling for multiple tests.
Area of Science:
- Epidemiology
- Biostatistics
- Health Informatics
Background:
- Longitudinal observational databases are crucial for public health surveillance.
- Tree-based scan statistics offer a method for data mining in large-scale health datasets.
- Adjusting for multiple testing is essential when analyzing numerous disease outcomes.
Purpose of the Study:
- To provide sample size calculations for implementing tree-based scan statistics in longitudinal observational databases.
- To evaluate the statistical power of both unconditional and conditional Poisson tree-based scan statistics.
- To guide researchers in designing effective data mining analyses for pharmacovigilance.
Main Methods:
- Utilized tree-based scan statistics on epidemiologic datasets with hierarchical outcome structures.
- Assessed statistical power by varying excess risk, sample size, event rates, and healthcare utilization.
- Quantified power reduction due to inappropriate risk window specifications.
Main Results:
- Achieved at least 98% power to detect an excess risk of 1 event per 10,000 exposed individuals with a sample size of 500,000.
- The conditional tree-based scan statistic effectively controlled Type I error in the presence of confounding healthcare utilization, unlike the unconditional version.
- Demonstrated the impact of risk window size on statistical power.
Conclusions:
- Tree-based scan statistics enhance pharmacovigilance by enabling monitoring of numerous outcomes with controlled multiple testing.
- Power evaluations are critical for optimizing the design and implementation of retrospective data mining studies.
- These statistical tools support robust surveillance in large longitudinal health databases.
More Related Videos
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
10:43Eye-tracking Technology and Data-mining Techniques used for a Behavioral Analysis of Adults engaged in Learning Processes
Published on: June 10, 2021
Related Concept Videos
Bioequivalence Data: Statistical Interpretation
Statistical Methods for Analyzing Epidemiological Data
Statistical Methods to Analyze Parametric Data: ANOVA
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
Statistical Software for Data Analysis and Clinical Trials
Statistical Significance
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...