Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Cluster Sampling Method01:20

Cluster Sampling Method

11.5K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.5K
Statistical Software for Data Analysis and Clinical Trials01:12

Statistical Software for Data Analysis and Clinical Trials

428
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
428
Comparing the Survival Analysis of Two or More Groups01:20

Comparing the Survival Analysis of Two or More Groups

92
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
92
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

244
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
244
Statistical Methods to Analyze Parametric Data: ANOVA01:12

Statistical Methods to Analyze Parametric Data: ANOVA

251
Analysis of Variance, or ANOVA, is a powerful statistical technique used to analyze parametric data, primarily in research and experimental studies. It's designed to compare the means of two or more groups, assisting researchers in identifying any significant differences between these group means. There are two main types of ANOVA based on the complexity of the analysis: one-way and two-way.
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
251
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

5.3K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
5.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Continuous vs Intermittent Postoperative Vital Sign Monitoring: A Cluster Randomized Crossover Trial.

JAMA network open·2026
Same author

Imputation and Missing Indicators for Handling Missing Longitudinal Data: Data Simulation Analysis Based on Electronic Health Record Data.

JMIR medical informatics·2025
Same author

A comparison of random forest variable selection methods for regression modeling of continuous outcomes.

Briefings in bioinformatics·2025
Same author

Gerontologic Biostatistics and Data Science: Aging Research in the Era of Big Data.

The journals of gerontology. Series A, Biological sciences and medical sciences·2024
Same author

Statins Do Not Significantly Affect Oxidative Nitrosative Stress Biomarkers in the PREVENT Randomized Clinical Trial.

Clinical cancer research : an official journal of the American Association for Cancer Research·2024
Same author

Performance of Cox regression models for composite time-to-event endpoints with component-wise censoring in randomized trials.

Clinical trials (London, England)·2023

Related Experiment Video

Updated: May 15, 2025

Global and Current Research Trends of Single-Cell Sequencing in Cancer: A Bibliometric and Visualization Study
07:49

Global and Current Research Trends of Single-Cell Sequencing in Cancer: A Bibliometric and Visualization Study

Published on: April 18, 2025

63

OpenClustered: an R package with a benchmark suite of clustered datasets for methodological evaluation and

Nathaniel Sean O'Connell1, Jaime Lynn Speiser2

  • 1Department of Biostatistics and Data Science, Wake Forest University School of Medicine, Winston-Salem, NC, 27157, USA.

BMC Medical Research Methodology
|April 11, 2025
PubMed
Summary

A new R package, OpenClustered, provides 19 open-source clustered datasets for methodology comparison. This resource offers empirical guidance, reducing bias compared to traditional data simulation studies for clustered data analysis.

Keywords:
BenchmarkingClustered dataModel evaluation

More Related Videos

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

6.9K
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.0K

Related Experiment Videos

Last Updated: May 15, 2025

Global and Current Research Trends of Single-Cell Sequencing in Cancer: A Bibliometric and Visualization Study
07:49

Global and Current Research Trends of Single-Cell Sequencing in Cancer: A Bibliometric and Visualization Study

Published on: April 18, 2025

63
Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

6.9K
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.0K

Area of Science:

  • Statistics
  • Data Science
  • Bioinformatics

Background:

  • Clustered data, common in epidemiology and social sciences, exhibit correlations within groups.
  • Existing data repositories lack comprehensive resources for clustered datasets.
  • Traditional simulation studies for methodology evaluation can introduce bias.

Purpose of the Study:

  • To develop an open-source data repository for clustered datasets.
  • To facilitate methodologic comparison and benchmarking studies.
  • To provide an alternative to potentially biased data simulation.

Main Methods:

  • Development of the R package 'OpenClustered'.
  • Inclusion of 19 diverse clustered datasets with binary outcomes.
  • Creation of tutorials for data manipulation and analysis.

Main Results:

  • The OpenClustered package offers 19 clustered datasets of varying sizes and compositions.
  • Tutorials demonstrate dataset filtering, summarization, and benchmarking.
  • A pilot study compared Frequentist and Bayesian generalized linear mixed models using the package.

Conclusions:

  • OpenClustered serves as a valuable resource for benchmarking studies with open-source clustered data.
  • The package promotes empirical methodologic guidance, enhancing research rigor.
  • Future plans include expanding the dataset collection and user submission functionality.