Initial Data Analysis for Cancer Registries: A Structured Framework and Demonstration Using Slovenian Cancer Registry
Maja Jurtela1,2, Lara Lusa3,4, Tina Žagar1
1Slovenian Cancer Registry, Institute of Oncology Ljubljana, 1000 Ljubljana, Slovenia.
Background/Objectives:
Initial data analysis (IDA) is essential for valid and reproducible statistical analyses, but existing IDA frameworks were primarily developed for single-study datasets. Cancer registries (CRs) are extensive data systems characterized by continuous updates, repeated data extraction, multiple analytical uses, and evolving classification systems, which create requirements not addressed by existing IDA frameworks. This study aims to develop a structured IDA framework adapted to CRs.
Methods:
We conceptualized IDA in CRs as a process spanning three data states: operational registry data, the extracted dataset and the analysis-ready dataset. The framework was developed by adapting existing IDA principles to the CR setting and organizing them into four stages: metadata, cleaning, screening, and reporting. The approach is based on predefined and versioned rule sets, structured recording of data processing, and metadata-based linkage between dataset definitions, data cleaning rules, screening outputs, the final report and dataset. A demonstrative survival dataset from the Slovenian Cancer Registry was used to illustrate implementation.
Results:
The framework is represented by the item set defining the activities and expected outputs of the structured IDA process in CRs. In the use case, metadata specified the dataset scope, intended use, variables, coding context, and applicable rules. Execution of the selected rules produced an analysis-ready dataset with traceable data cleaning steps, documented eligibility decisions, screening outputs describing the general and analysis-specific data properties, and the IDA report intended for external researchers to be delivered alongside the data.
Conclusions:
The proposed framework extends existing IDA approaches to meet the specific requirements of CRs. It supports consistent dataset preparation and with that improves transparent and reproducible data use.
Related Concept Videos
Cancer Survival Analysis
Statistical Methods for Analyzing Epidemiological Data
Statistical Software for Data Analysis and Clinical Trials
Comparing the Survival Analysis of Two or More Groups
Study Designs in Epidemiology
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and case-control studies.
Introduction To Survival Analysis
The primary goal of survival analysis is to estimate survival time—the time until a...

