Related Experiment Video
Updated: Aug 1, 2025

12:11
Simultaneous Affinity Enrichment of Two Post-Translational Modifications for Quantification and Site Localization
Published on: February 27, 2020
6.9K
iEnhancer-ELM: improve enhancer identification by extracting position-related multiscale contextual information based
Jiahao Li1, Zhourun Wu1, Wenhao Lin1
1School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen, Guangdong 518055, China.
Bioinformatics Advances
|April 28, 2023
Summary
A new method, iEnhancer-ELM, uses BERT-like language models to identify enhancers by learning position-related sequence information. This approach outperforms existing methods and aids in discovering enhancer motifs and biological mechanisms.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Enhancers are crucial cis-regulatory elements controlling gene transcription.
- Existing feature extraction methods struggle to capture position-dependent, multi-scale contextual information from DNA sequences.
Purpose of the Study:
- To propose a novel enhancer identification method, iEnhancer-ELM, leveraging BERT-like enhancer language models.
- To address the limitations of current methods in learning position-related multiscale contextual information from raw DNA sequences.
Main Methods:
- iEnhancer-ELM tokenizes DNA sequences using multi-scale k-mers.
- It employs a multi-head attention mechanism to extract positional and contextual information from different k-mer scales.
- The method evaluates and ensembles various k-mer scales for optimal performance.
Main Results:
- The proposed iEnhancer-ELM model demonstrates superior performance compared to state-of-the-art methods on benchmark datasets.
- A case study revealed 30 potential enhancer motifs using a 3-mer model, with 12 validated by STREME and JASPAR.
- The model shows interpretability and potential for uncovering enhancer biological mechanisms.
Conclusions:
- iEnhancer-ELM offers an effective approach for enhancer identification by capturing essential sequence features.
- The method provides insights into enhancer motifs and their regulatory roles.
- The developed models and code are publicly available for further research.
Related Concept Videos
Position-effect Variegation
6.4K
In 1928, a German botanist Emil Heitz observed the moss nuclei with a DNA binding dye. He observed that while some chromatin regions decondense and spread out in the interphase nucleus, others do not. He termed them euchromatin and heterochromatin, respectively. He proposed that the heterochromatin regions reflect a functionally inactive state of the genome. It was later confirmed that heterochromatin is transcriptionally repressed, and euchromatin is transcriptionally active chromatin.
6.4K
Conserved Binding Sites
4.3K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.3K
Multi-species Conserved Sequences
4.0K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.0K

