Related Experiment Video
Updated: May 1, 2026

gP2S, an Information Management System for CryoEM Experiments
Published on: June 10, 2021
The eGenVar data management system--cataloguing and sharing sensitive data and metadata for the life sciences
Sabry Razick1, Rok Močnik, Laurent F Thomas
1Department of Cancer Research and Molecular Medicine, Norwegian University of Science and Technology, Prinsesse Kristinasgt. 1, NO-7491 Trondheim, Norway and Department of Computer and Information Science, Norwegian University of Science and Technology, Sem Sælands vei 9, NO-7491 Trondheim, Norway.
This study introduces a metadata cataloguing system and software suite for life sciences data. It enables efficient data discovery and sharing while respecting privacy by keeping sensitive data in original locations.
Area of Science:
- Life Sciences
- Bioinformatics
- Data Management
Background:
- Systematic data management and controlled sharing are crucial for research reproducibility and efficiency.
- Storing sensitive data (e.g., from biobanks, clinical studies) in public repositories is often restricted due to legal and privacy concerns.
Purpose of the Study:
- To describe a metadata cataloguing system and software suite for managing and sharing life sciences data.
- To provide a framework for cataloguing both public and private data without centralizing sensitive information.
Main Methods:
- Developed a system storing three types of metadata: file information, file provenance/data lineage, and content descriptions.
- Created a software suite with graphical and command-line interfaces for users to report and tag files with metadata.
- Ensured original data files remain in their existing locations with access controls.
Main Results:
- The system successfully catalogues metadata for life sciences data.
- Users can efficiently report and tag files, enhancing data discoverability.
- The framework supports the cataloguing and sharing of both public and private data.
Conclusions:
- The metadata cataloguing system and software suite offer a common framework for managing and sharing diverse life sciences data.
- This approach facilitates data reproducibility and reduces redundancy while maintaining data privacy and security.
- The system enables efficient location of complementing or contradicting information across distributed datasets.

