Related Experiment Video
Updated: Dec 15, 2025

Executing Complexity-Increasing Queries in Relational MySQL and NoSQL MongoDB and EXist Size-Growing ISO/EN 13606 Standardized EHR Databases
Published on: March 19, 2018
Enabling ad-hoc reuse of private data repositories through schema extraction
Lars Christoph Gleim1, Md Rezaul Karim2,3, Lukas Zimmermann4
1Informatik 5, RWTH Aachen University, Ahornstr. 55, Aachen, 52062, Germany. gleim@cs.rwth-aachen.de.
This study introduces an automated schema extraction method to query sensitive data without direct access, enabling secure data reuse. This approach facilitates data integration and analysis, crucial for advancing data-driven biomedical research.
Area of Science:
- Biomedical Informatics
- Data Science
- Computer Science
Background:
- Legal and ethical restrictions, like the EU General Data Protection Rules (GDPR), limit sharing sensitive data across organizations.
- The Personal Health Train initiative explores decentralized data utilization to bypass data transfer needs.
- Existing systems face challenges in accessing and integrating sensitive data stored in disparate repositories.
Purpose of the Study:
- To propose a configurable and automated approach for schema extraction and publishing.
- To enable ad-hoc SPARQL query formulation against RDF triple stores without direct data access.
- To facilitate secure query execution under data provider control.
Main Methods:
- Developed a configurable and automated system for schema extraction and publishing.
- Ensured compatibility with existing Semantic Web technologies.
- Enabled schema introspection-assisted authoring of SPARQL queries.
Main Results:
- Successfully derived a configurable amount of concise, task-relevant schema for four distinct datasets.
- Demonstrated the ability to formulate SPARQL queries against RDF triple stores without direct access to private data.
- Validated the approach for enabling schema introspection-assisted query authoring.
Conclusions:
- Automated schema extraction and publishing supports introspection-assisted query creation for data selection and integration.
- The system architecture enables data reuse from private repositories where shared schemas are infeasible.
- This approach advances the reuse of data from previously inaccessible sources, promoting data-driven methods in biomedicine.
More Related Videos
Related Concept Videos
Impact of Schemas
Schemas
Self-Schemas
Storage
Schemata
Two types of schemata are:
Extraction: Advanced Methods

