Related Experiment Video
Updated: Jun 29, 2025

14:49
Associated Chromosome Trap for Identifying Long-range DNA Interactions
Published on: April 23, 2011
14.5K
Joint Representation Learning for Retrieval and Annotation of Genomic Interval Sets
Erfaneh Gharavi1,2, Nathan J LeRoy1,3, Guangtao Zheng4
1Center for Public Health Genomics, School of Medicine, University of Virginia, Charlottesville, VA 22908, USA.
Bioengineering (Basel, Switzerland)
|March 27, 2024
Summary
This study introduces a novel representation learning system for fast and flexible searching of large genomic interval databases. The method uses co-embeddings to improve information retrieval accuracy for genomic region data.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomic Data Analysis
Background:
- Genomic interval data is rapidly increasing, necessitating efficient search systems.
- Current methods like string matching and region overlap analysis have limitations in accuracy and computational cost.
- There is a need for advanced methods to query large, complex genomic interval databases.
Purpose of the Study:
- To develop a novel genomic interval search system using representation learning.
- To capture similarities between genomic region sets and their metadata in a low-dimensional space.
- To enable fast, flexible, and accurate information retrieval for genomic region data.
Main Methods:
- Developed a system that trains numerical embeddings for region sets and metadata simultaneously.
- Utilized representation learning to create co-embeddings in a low-dimensional space.
- Employed embedding distance computations for information retrieval tasks.
Main Results:
- The system successfully performs three key information retrieval tasks: query string matching, label suggestion, and similar region set retrieval.
- Jointly learned representations of region sets and metadata demonstrate effectiveness.
- Achieved fast, flexible, and accurate genomic region information retrieval.
Conclusions:
- Representation learning offers a promising approach for enhancing genomic interval database searches.
- The developed system addresses limitations of existing methods for querying large-scale genomic data.
- This method facilitates more efficient and accurate analysis of genomic region information.
Related Concept Videos
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Gene Duplication and Divergence
6.1K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
6.1K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
Genomics
36.3K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.3K
Genomic Imprinting and Inheritance
34.4K
Diploid organisms inherit genetic material through chromosomes from both parents. Copies of the same gene are known as alleles. In most cases, both alleles are simultaneously expressed and allow various cellular processes to function optimally. If one of the alleles is missing or mutated, the expression of the other allele can compensate; however, this is not true for all genes.
The expression of some genes depends on which parent passed the gene to the offspring, through a phenomenon known as...
The expression of some genes depends on which parent passed the gene to the offspring, through a phenomenon known as...
34.4K
DNA Microarrays
17.4K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.4K

