Related Experiment Videos
HGVbase: a human sequence variation database emphasizing data quality and a broad spectrum of data sources.
D Fredman1, M Siegfried, Y P Yuan
1Center for Genomics and Bioinformatics, Karolinska Institute, Berzelius väg, S171 77 Stockholm, Sweden.
Nucleic Acids Research
|December 26, 2001
Summary
HGVbase provides a comprehensive, non-redundant database of human genomic variations, including single nucleotide polymorphisms (SNPs) and mutations. It offers advanced search tools and data downloads to aid genetic research and disease mutation analysis.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- Genomic variation data is rapidly expanding, requiring curated and non-redundant databases.
- Accurate and accessible variation data is crucial for understanding human genetics and disease.
- Existing databases may lack comprehensive features or suffer from redundancy.
Purpose of the Study:
- To establish HGVbase as a high-quality, non-redundant repository for human genomic variation data.
- To facilitate data interrogation through advanced search functionalities.
- To support research by providing diverse data formats and integrated information.
Main Methods:
- Curating and integrating various types of genomic variations, primarily single nucleotide polymorphisms (SNPs).
- Implementing online search tools for sequence similarity, keyword queries, and genome coordinates.
- Providing data in multiple formats (XML, Fasta, SRS, SQL, tagged-text) for broad accessibility.
- Ensuring data quality through semi-automated checking and automated annotation processes.
- Mapping variants to the draft genome sequence and referencing EMBL/GenBank files.
Main Results:
- HGVbase offers a centralized, non-redundant collection of human genomic variations, including neutral polymorphisms and disease-related mutations.
- Search functionalities allow efficient data retrieval by sequence, keywords, and genomic location.
- Data is presented with surrounding sequence context, gene information, and population allele frequencies where available.
- Automated annotation and data checking ensure internal consistency and accuracy.
- Extended data structures support haplotype and genotype information capture.
Conclusions:
- HGVbase serves as a valuable resource for genomic variation data, enhancing research accessibility and utility.
- The database supports the study of both common genetic variations and rare disease mutations.
- Ongoing initiatives aim to expand the repository to include a comprehensive collection of clinical mutations and associated phenotypes.