Related Experiment Video
Updated: May 28, 2025

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
Enhancing clinical data warehousing with provenance data to support longitudinal analyses and large file management:
Maxime Wack1, Adrien Coulet2, Anita Burgun3
1Centre de Recherche des Cordeliers, UMRS 1138, Inserm, Université Paris Cité, Sorbonne Université, Paris, France; Inria Paris, Paris, France; Department of Biomedical Informatics, Hôpital Européen Georges Pompidou, AP-HP, Paris, France; Centre Hospitalier National d'Ophtalmologie des Quinze-Vingts, IHU FOReSIGHT, 75012 Paris, France.
Background:
If hospital Clinical Data Warehouses are to address today's focus in personalized medicine, they need to be able to track patients longitudinally and manage the large data sets generated by whole genome sequencing, RNA analyses, and complex imaging studies. Current Clinical Data Warehouses address neither issue. This paper reports on methods to enrich current systems by providing provenance data allowing patient histories to be followed longitudinally and managing the linking and versioning of large data sets from whatever source. The methods are open source and applicable to any clinical data warehouse system, whether data schema it uses.
Method:
We introduce gitOmmix, an approach that overcomes these limitations, and illustrate its usefulness in the management of medical omics data. gitOmmix relies on (i) a file versioning system: git, (ii) an extension that handles large files: git-annex, (iii) a provenance knowledge graph: PROV-O, and (iv) an alignment between the git versioning information and the provenance knowledge graph.
Results:
Capabilities inherited from git and git-annex enable retracing the history of a clinical interpretation back to the patient sample, through supporting data and analyses. In addition, the provenance knowledge graph, aligned with the git versioning information, enables querying and browsing provenance relationships between these elements.
Conclusion:
gitOmmix adds a provenance layer to CDWs, while scaling to large files and being agnostic of the CDW system. For these reasons, we think that it is a viable and generalizable solution for omics clinical studies.
More Related Videos
09:52A Clinical Metaproteomics Workflow Implemented within Galaxy Bioinformatics Platform to Analyze Host-Microbiome Interactions Underlying Human Disease
Published on: January 10, 2025
11:18Generation of Comprehensive Thoracic Oncology Database - Tool for Translational Research
Published on: January 22, 2011
Related Concept Videos
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Genomics