Related Experiment Video
Updated: Sep 7, 2026

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
Published on: November 22, 2019
Assessing the fidelity of synthetic relational databases with high-dimensional categorical data: Proposal of a
Roxane Girault1, Jassim Bensafir2, Antoine Lamer3
1Univ. Lille, CHU Lille, ULR 2694 METRICS, Cerim, F-59000 Lille, France.
Background:
Although data reuse is increasingly common in healthcare research, changing regulatory frameworks are impeding these efforts. Synthetic data constitute a promising issue but also create challenges, such as the assessment of data fidelity and the ability to output identical statistical results. Conventional validation approaches often rely on univariate or bivariate comparisons, which fail to capture the complexity of associations between multivalued categorical variables.
Material:
We built and studied a fictitious database of 10,000 hospital stays reproducing the structure of the French Programme de Médicalisation des Systèmes d'Information database. Each stay included single-valued variables (one value per individual: sex, age in deciles, and diagnosis-related group) and multivalued variables (zero, one or several values per individual: diagnoses coded according to the International Classification of Diseases, 10th Edition, and procedures coded according to the French Classification Commune des Actes Médicaux).
Method:
All categorical variables were binarized, and thousands of pairwise association metrics (primarily odds ratios) were calculated for the reference and evaluation datasets. The results were summarized using curves, bubble charts, heatmaps, and coefficients such as exponential mean deviation. Simulated data degradations from 0% to 100% were introduced to evaluate the method's sensitivity.
Results:
We analyzed 500 ICD-10 diagnoses and 450 CCAM procedures, representing 225,000 possible combinations. We developed and evaluated graphical representations for assessing data fidelity at a glance. In simulations of an increasing degree of data degradation, those graphical representations and comprehensive, quantitative metrics facilitated the detection of the gradual loss of data fidelity.
Conclusion:
We developed a simple, scalable, agnostic framework for assessing the fidelity of healthcare databases by systematically analyzing associations among the modalities of coded variables. This method complements existing approaches. It is particularly suitable for the evaluation of synthetic relational databases because it offers both general and granular insights into data fidelity loss.
More Related Videos
Related Concept Videos
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Impact of Schemas
Dimensional Analysis
In fluid mechanics, dimensional...
Dimensional Analysis
Conversion Factors and Dimensional Analysis
The unit...
Dimensional Analysis
Dimensional analysis allows us to analyze and compare physical quantities on a...
Dimensional Analysis

