Related Experiment Videos
Gene maps and location databases
1Department of Community Medicine, University of Southampton, U.K.
Annals of Human Genetics
|July 1, 1991
Summary
A novel location database organizes genetic and physical locus data in linear space for efficient genome analysis. This approach offers a scalable model for human genome database development, improving upon traditional interval databases.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Traditional interval databases define genomic locations in metric space, requiring complex list-processing for data retrieval.
- Existing interval database structures present challenges due to overlapping intervals and varying ordinal assignments across tables.
- Well-studied experimental organisms have successfully utilized location databases for genetic and physical mapping.
Purpose of the Study:
- To define and describe a location database model for organizing genomic data.
- To contrast the proposed location database with existing interval database structures.
- To establish principles for developing a human genome database based on prior organismal studies.
Main Methods:
- Defining a location database using a vector of genetic and physical locations for each locus.
- Employing virtual sorting on composite location for ordering loci within the database.
- Comparing the linear space organization of the location database with the metric space of interval databases.
Main Results:
- The location database organizes loci in linear space via ordered vectors of genetic and physical locations.
- This contrasts with interval databases where location inference is complex and data is fragmented.
- The location database model has proven effective in numerous well-studied experimental organisms.
Conclusions:
- A location database provides a robust framework for organizing complex genomic data.
- Principles derived from existing location databases can guide the development of a comprehensive human genome database.
- The linear space model offers a more efficient and scalable approach to genomic data management.