Initial Data Analysis for Cancer Registries: A Structured Framework and Demonstration Using Slovenian Cancer Registry
Maja Jurtela1,2, Lara Lusa3,4, Tina Žagar1
1Slovenian Cancer Registry, Institute of Oncology Ljubljana, 1000 Ljubljana, Slovenia.
Cancers
|July 28, 2026
Summary
This study introduces a new Initial Data Analysis (IDA) framework specifically designed for cancer registries (CRs). The framework enhances data transparency and reproducibility for researchers using CR data.
Area of Science:
- Biostatistics
- Data Management
- Cancer Epidemiology
Background:
- Existing Initial Data Analysis (IDA) frameworks are not optimized for the dynamic nature of cancer registries (CRs).
- CRs present unique challenges due to continuous updates, multiple data uses, and evolving classification systems.
- A tailored IDA framework is needed to ensure valid and reproducible statistical analyses from CR data.
Purpose of the Study:
- To develop and present a structured Initial Data Analysis (IDA) framework specifically adapted for cancer registries (CRs).
- To address the limitations of existing IDA frameworks in the context of complex CR data environments.
- To enhance the transparency and reproducibility of data preparation and analysis for CR datasets.
Main Methods:
- Conceptualized IDA across three data states: operational registry, extracted, and analysis-ready datasets.
- Adapted existing IDA principles into a four-stage framework: metadata, cleaning, screening, and reporting.
- Employed predefined, versioned rule sets, structured data processing records, and metadata linkage for traceability.
Main Results:
- The developed framework outlines specific activities and outputs for structured IDA in CRs.
- Metadata captured dataset scope, intended use, variables, and coding context.
- Implementation demonstrated traceable data cleaning, documented decisions, screening outputs, and a comprehensive IDA report.
Conclusions:
- The proposed framework extends current IDA methodologies to meet the specific demands of cancer registries.
- It facilitates consistent dataset preparation, leading to improved transparency in data utilization.
- The framework ultimately supports more reproducible research using cancer registry data.
Related Concept Videos
Cancer Survival Analysis
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
Statistical Methods for Analyzing Epidemiological Data
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
Statistical Software for Data Analysis and Clinical Trials
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
Comparing the Survival Analysis of Two or More Groups
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and Cox...
Study Designs in Epidemiology
Epidemiological study designs are fundamental tools for investigating the distribution, determinants, and control of health conditions in populations. They help researchers understand the relationships between exposures and outcomes, and they broadly fall into two categories: "observational" and "experimental" studies.
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and case-control studies.
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and case-control studies.
Introduction To Survival Analysis
Survival analysis is a statistical method used to study time-to-event data, where the "event" might represent outcomes like death, disease relapse, system failure, or recovery. A unique feature of survival data is censoring, which occurs when the event of interest has not been observed for some individuals during the study period. This requires specialized techniques to handle incomplete data effectively.
The primary goal of survival analysis is to estimate survival time—the time until a...
The primary goal of survival analysis is to estimate survival time—the time until a...

