Systematically linking tranSMART, Galaxy and EGA for reusing human translational research data
Chao Zhang1, Jochem Bijlard1,2, Christine Staiger3
1Department of Computer Science, Vrije Universiteit Amsterdam, Amsterdam, 1081 HV, Netherlands.
F1000Research
|November 11, 2017
Summary
A new data ecosystem links raw and interpreted molecular profiling data using the European Genome-phenome Archive (EGA) and tranSMART. This enables efficient data management and reanalysis, promoting data reuse through FAIR persistent identifiers.
Area of Science:
- Bioinformatics
- Data Science
- Genomics
Background:
- High-throughput molecular profiling generates complex data requiring advanced computational interpretation.
- Explosive data growth necessitates robust data management for efficient organization and integration.
- Existing systems often lack seamless integration between raw and interpreted data.
Purpose of the Study:
- To design and implement a data ecosystem for linking raw and interpreted molecular profiling data.
- To establish a framework for efficient data management and computational workflow execution.
- To facilitate data discovery, traceability, and reanalysis for clinical research.
Main Methods:
- Established an ELIXIR implementation study in collaboration with the Translational research IT (TraIT) programme.
- Utilized the European Genome-phenome Archive (EGA) for raw data storage.
- Employed tranSMART for interpreted and clinical data, and Galaxy for computational workflows.
- Integrated data by systematically linking repositories and using FAIR persistent identifiers.
Main Results:
- Developed a data ecosystem integrating EGA, tranSMART, and Galaxy.
- Successfully structured TraIT Cell Line Use Case (TraIT-CLUC) data for storage and cross-referencing.
- Enabled data flow from EGA to Galaxy for raw data reanalysis.
- Demonstrated user ability to select cohorts in tranSMART, trace to raw data, and perform reanalysis in Galaxy.
Conclusions:
- The proposed data ecosystem effectively links raw and interpreted molecular profiling data.
- FAIR persistent identifiers are crucial for stable data linkage and reuse across different data ontology levels.
- This approach enhances data discoverability, traceability, and reanalysis capabilities in clinical research.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
15.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.2K
Improving Translational Accuracy
3.7K
3.7K
Evolutionary Relationships through Genome Comparisons
7.1K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
7.1K


