LOGOWheat: deep learning-based prediction of regulatory effects for noncoding variants in wheats
Lingpeng Kong1, Hong Cheng1, Kun Zhu1,2
1Agricultural Genomics Institute at Shenzhen, Chinese Academy of Agricultural Sciences, No. 97 Buxin Road, Dapeng New District, Shenzhen 518124, China.
Briefings in Bioinformatics
|January 10, 2025
Summary
We developed Language of Genome for Wheat (LOGOWheat), a deep learning tool to predict regulatory effects of noncoding variants in wheat. This method accurately identifies functional genetic variations, aiding crop improvement.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Identifying regulatory effects of noncoding genetic variants is crucial but challenging.
- Advancements in wheat epigenomic data offer new avenues for functional variant modeling.
Purpose of the Study:
- To introduce Language of Genome for Wheat (LOGOWheat), a deep learning tool for predicting regulatory effects of noncoding variants in wheat.
- To enable accurate identification and prioritization of functional genetic variations in wheat.
Main Methods:
- Utilized a self-attention-based, contextualized pretrained language model for wheat genome representation.
- Fine-tuned the model with epigenomic profiling data to discern regulatory codes in genomic sequences.
- Employed deep learning to predict chromatin features and variant impacts.
Main Results:
- LOGOWheat achieved high accuracy in predicting chromatin features, with an average AUROC of 0.8531 and AUPRC of 0.7633.
- Demonstrated LOGOWheat's utility in scoring and prioritizing causal variants and in silico mutagenesis mapping.
- Showcased the tool's ability to identify high-impact sites and functional motifs.
Conclusions:
- LOGOWheat is an effective deep learning tool for predicting regulatory effects of noncoding variants in wheat.
- The tool facilitates the discovery of functional variations and aids in understanding wheat genome regulation.
- Integration with evolutionary conservation information can enhance the extraction of potential functional variations from wheat populations.
Related Concept Videos
lncRNA - Long Non-coding RNAs
8.5K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
8.5K
Cis-regulatory Sequences
9.7K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
9.7K
Regulation of Expression at Multiple Steps
868
The gene expression in cells is regulated at different stages: (i) transcription, (ii) RNA processing, (iii) RNA localization, and (iv) translation. Transcriptional regulation is mediated by regulatory proteins such as transcription factors, activators, or repressors—these control gene expression by initiating or inhibiting the transcription of genes. Once a precursor or pre-mRNA is produced, it undergoes post-transcriptional modification, including 5' capping, splicing, and the...
868
Genome-wide Association Studies-GWAS
12.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.4K


