Related Experiment Video
Updated: Sep 24, 2025

09:32
Deciphering High-Resolution 3D Chromatin Organization via Capture Hi-C
Published on: October 14, 2022
3.7K
Integrating convolution and self-attention improves language model of human genome for interpreting non-coding
Meng Yang1,2, Lichao Huang1, Haiping Huang1
1MGI, BGI-Shenzhen, Shenzhen 518083, China.
Nucleic Acids Research
|May 10, 2022
Summary
We introduce LOGO (Language of Genome), a lightweight deep learning model that interprets the non-coding genome. It improves predictions of gene regulatory elements and prioritizes disease-associated variants.
Area of Science:
- Genomics
- Computational Biology
- Bioinformatics
Background:
- Interpreting the non-coding genome is crucial for human genetics but challenging due to the vastness of biochemical elements.
- Deep learning offers promising computational approaches for non-coding region interpretation.
Purpose of the Study:
- To develop a computationally efficient deep learning model, LOGO (Language of Genome), for interpreting the human non-coding genome.
- To enhance sequence labeling and variant prioritization tasks using self-supervised learning on the human reference genome.
Main Methods:
- Developed LOGO, a self-attention based contextualized pre-trained language model with a light architecture (1 million parameters).
- Employed self-supervision for bidirectional genome representation learning.
- Fine-tuned LOGO for sequence labeling and extended it for variant prioritization using convolutional modules and special input encoding.
Main Results:
- Achieved 15% absolute improvement in promoter identification and 4.5% in enhancer-promoter interaction prediction.
- Demonstrated state-of-the-art multi-task predictive power on thousands of chromatin features with significantly fewer parameters than existing models.
- Improved sensitivity and specificity in prioritizing non-coding variants associated with human diseases.
Conclusions:
- LOGO provides an accurate, fast, scalable, and robust framework for interpreting non-coding genomic regions.
- The model successfully inferred regulatory mechanisms for type 2 diabetes GWAS signals.
- LOGO advances base-resolution interpretation of the non-coding genome for sequence labeling and variant prioritization.
Related Concept Videos
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Genome Annotation and Assembly
19.4K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.4K
RNA-seq
10.4K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.4K
Super-resolution Fluorescence Microscopy
8.1K
Super-resolution fluorescence microscopy (SRFM) provides a better resolution than conventional fluorescence microscopy by reducing the point spread function (PSF). PSF is the light intensity distribution from a point that causes it to appear blurred. Due to PSF, each fluorescing point appears bigger than its actual size, and it is the PSF interference of nearby fluorophores that causes the blurred image. Various approaches to achieving higher resolution through SRFM have recently been...
8.1K
Genome-wide Association Studies-GWAS
14.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.5K
Genome Size and the Evolution of New Genes
2.7K
2.7K

