Related Experiment Video
Updated: Aug 31, 2025

Investigating Protein Sequence-structure-dynamics Relationships with Bio3D-web
Published on: July 16, 2017
Bio-Strings: A Relational Database Data-Type for Dealing with Large Biosequences
Sergio Lifschitz1, Edward H Haeusler1, Marcos Catanho2
1Departamento de Informática, Pontifícia Universidade Católica do Rio de Janeiro (PUC-Rio), Rio de Janeiro 22451-900, Brazil.
Relational databases can effectively store and manage long biological data strings from DNA sequencers. This approach ensures data integrity and efficient handling of variable-sized genomic sequences.
Area of Science:
- Bioinformatics
- Database Management Systems
- Genomics
Background:
- DNA sequencers generate extensive biological data strings.
- Current text file systems face challenges in storing and efficiently managing large genomic datasets.
- Handling variable-sized biological sequences while preserving their meaning is crucial.
Purpose of the Study:
- To propose a novel approach for representing and manipulating biological sequences within relational databases.
- To explore the suitability of relational text data types for genomic data storage.
- To develop and evaluate a logical schema and associated functions for genomic data management.
Main Methods:
- Implementation of a relational text data type within a relational database management system (RDBMS).
- Development of a logical schema for core biological information representation.
- Evaluation of stored functions for efficacy and efficiency in handling genomic data.
Main Results:
- Demonstrated the feasibility of enforcing basic and complex genomic data requirements using the proposed relational text data type.
- Showcased efficient storage and manipulation of variable-sized biological sequences.
- Validated the effectiveness of the implemented logical schema and functions.
Conclusions:
- Relational database management systems (RDBMS) with relational text data types can effectively manage and persist biological sequences.
- The proposed domain-specific abstract data type approach simplifies data handling for non-technical users.
- This method offers a robust solution for genomic data storage and analysis.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Single-Strand DNA Binding Proteins
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Sanger Sequencing

