Related Experiment Video
Updated: May 21, 2026

A User-friendly and Powerful R Analysis of Large-scale Datasets
Published on: November 4, 2025
Understanding the limits of large datasets
Catherine M Sanders1, Sidney L Saltzstein, Matthew M Schultzel
1Rebecca and John Moores UCSD Cancer Center, University of California San Diego, La Jolla, CA 92093-0850, USA. Catherine.Sanders@osumc.edu
Abstract:
Many health professionals use large datasets to answer behavioral, translational, or clinical questions. Understanding the impact of missing data in large databases, such as disease registries, can avoid erroneous interpretations of these data. Using the California Cancer Registry, the authors selected seven common cancers, seven sociodemographic and clinical variables, and the top three reporting sources, as examples of the type of data that would be deemed critical to most studies. The gender variable had no missing data, followed by age (<0.1 % missing), ethnicity (1.7 %), stage (9.8 %), differentiation (39.1 %), and birthplace (41.1 %). Reports from hospitals and clinics had the lowest percentages of missing data. Users of large datasets should anticipate the limitations of missing data to prevent methodological flaws and misinterpretations of research findings. Knowledge of what and how much data may be missing in large datasets can help prevent errors in research conclusions, while better guiding treatment modalities and public health policies and programs.
Related Concept Videos
Central Limit Theorem
The sample size, n, that...
Limits at Infinity
Introduction to Limits
The Precise Definition of a Limit
Maximum Size of Aggregate
Types of Limits I

