Related Experiment Video
Updated: Jan 12, 2026

TBase - an Integrated Electronic Health Record and Research Database for Kidney Transplant Recipients
Published on: April 13, 2021
Attention-based synthetic data generation for calibration-enhanced survival analysis: A case study for chronic kidney
Nicholas I-Hsien Kuo1, Blanca Gallego1, Louisa Jorm1
1Centre for Big Data Research in Health (CBDRH), The University of New South Wales, Sydney 2033, New South Wales, Australia.
Objectives:
Access to real-world healthcare data is constrained by privacy regulations and data imbalances, hindering the development of fair and reliable clinical prediction models. Synthetic data offers a potential solution, yet existing methods often fail to maintain calibration or enable subgroup-specific augmentation. This study introduces Masked Clinical Modelling (MCM), an attention-based synthetic data generation framework designed to enhance survival model calibration in both global and stratified analyses.
Methods:
MCM uses masked feature reconstruction to learn feature dependencies without explicitly training on survival objectives. It supports both standalone dataset synthesis and conditional data augmentation, enabling the generation of targeted synthetic subcohorts without retraining. Evaluated on a chronic kidney disease (CKD) electronic health record (EHR) dataset, MCM was benchmarked against eight baseline methods, including variational autoencoders, GANs, SMOTE variants, and a recent risk-aware distillation model. Model performance was assessed via calibration loss, Cox model consistency, and Kaplan-Meier fidelity.
Results:
MCM-generated data closely replicated statistical properties of the real dataset, pre- served hazard ratios, and matched time-to-event curves with high fidelity. Cox models trained on MCM-augmented data demonstrated improved calibration, reducing overall calibration loss by 15% and subgroup meta-calibration loss by 9% compared to unaugmented data. These improvements held across multiple high-risk subgroups including those with diabetes, renal dys- function, and advanced age. Unlike competing methods, MCM achieved this without retraining or outcome-specific tuning.
Conclusions:
MCM offers a practical and flexible framework for generating synthetic survival data that improves risk model calibration. By supporting both reproducible dataset synthesis and conditional subgroup augmentation, MCM bridges privacy-preserving data access with calibration-aware learning. This work highlights the role of synthetic data not just as a privacy tool, but as a vehicle for improving equity and reliability in clinical modelling.
Related Concept Videos
Drug Dosing in Renal Diseases: Estimation of Glomerular Filtration Rate Based on Serum Creatinine Concentration
Chronic Kidney Disease III: Interprofessional Care
Chronic Kidney Disease I: Introduction
Chronic Kidney Disease IV: Nursing Management
Chronic Kidney Disease II: Clinical Manifestations
Acute Kidney Injury IV: Diagnostic Studies and Prevention

