预培训基因组语言模型与更好的功能基因组模型的变体
bioRxiv : the preprint server for biology
|September 2, 2025
概括
我们开发了UKBioBERT, 一种DNA语言模型, 这种方法提高了对基因调节和基因变异效应的理解.
科学领域:
- 基因组学
- 计算生物学
- 生物信息学
背景情况:
- 基因组语言模型 (GLM) 从DNA序列中学习以表示基因组上下文.
- 序列到功能 (S2F) 模型将遗传信息与基因表达和表型联系起来.
- 为了个性化基因表达预测,将GLM和S2F模型连接起来仍然是一个挑战.
研究的目的:
- 使用英国生物银行遗传数据开发一种新的DNA语言模型 - - UKBioBERT.
- 将UKBioBERT与现有的S2F模型 (Enformer,Borzoi) 集成,以创建增强的预测模型.
- 改善基因表达水平的预测,了解基因变异的影响.
主要方法:
- 在英国生物银行对基因变异进行DNA语言模型 (UKBioBERT) 的预训练.
- 通过UKBioBERT生成信息序列嵌入.
- 将UKBioBERT嵌入式与S2F架构 (Enformer,Borzoi) 结合起来,形成UKBioFormer和UKBioZoi.
- 评估不同群体基因表达预测的模型性能.
主要成果:
- UKBioBERT嵌入有效地识别基因功能并改善细胞系中的基因表达预测.
- 在预测高度可预测的基因表达水平方面,UKBioFormer和UKBioZoi表现出卓越的性能.
- 综合模型在多样化的群体中很好地泛化.
- UKBioFormer准确地捕捉了基因型-表型关系,使得基基因突变分析成为可能.
结论:
- 将基因组语言模型与序列对函数的方法集成,可以显著地推进功能基因组学.
- UKBioBERT为了解基因功能和表达的可预测性提供了有价值的嵌入.
- 开发的UKBioFormer和UKBioZoi模型为预测基因表达和分析遗传变异效应提供了改进的工具.
更多相关视频
08:04Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
492
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
34.0K
相关概念视频
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
Genomics
37.4K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
37.4K
Genetic Lingo
104.5K
Overview
104.5K
Genome Annotation and Assembly
19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K
Genome-wide Association Studies-GWAS
14.1K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.1K
Comparing Copy Number Variations and SNPs
17.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.9K
