用GeneTEA对基因描述进行自然语言处理,以进行过度代表性分析
Isabella A Boyle1, Nayeem Akram Aquib2, Mustafa Kocak2
1Broad Institute of MIT and Harvard, Cambridge, MA, 02142, USA. iboyle@broadinstitute.org.
Genome biology
|October 31, 2025
概括
GeneTEA使用基因描述的自然语言处理来创建一个新的基因组数据库. 这种工具可以准确地识别生物丰富,比现有方法更少的错误发现和更少的冗余.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 过度代表性分析是识别基因列表中的生物丰富的一种常见方法.
- 现有的工具往往难以控制错误发现率,并可能产生冗余的结果.
研究的目的:
- 介绍GeneTEA,一种用于基因组丰富分析的新型模型.
- 开发一个新的基因组数据库,使用对基因描述的自然语言处理.
- 提高确定生物丰富的准确性和减少冗余性.
主要方法:
- 基因TEA吸收了自由文本基因描述.
- 它采用自然语言处理方法来学习稀疏的基因逐术语嵌入.
- 该模型的性能与现有的过度代表性分析工具进行了比较.
主要成果:
- 基因TEA有效控制了错误发现率.
- 该模型始终确定了最相关的生物见解.
- 与传统方法相比,GeneTEA 以减少冗余的方法实现了这一目标.
- 这种方法可以适应其他生物和化合物.
结论:
- 基因TEA提供了一种强大而准确的方法,用于过度代表性分析.
- 该模型提供了一个由基因描述衍生的新基因集合数据库.
- GeneTEA通过控制错误发现和尽量减少冗余来改进现有的工具.
- 一个交互式应用程序和API可用于GeneTEA模型.
更多相关视频
09:35A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017
18.3K
10:40Comprehensive Workflow for the Genome-wide Identification and Expression Meta-analysis of the ATL E3 Ubiquitin Ligase Gene Family in Grapevine
Published on: December 22, 2017
10.9K
相关概念视频
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K
Genome-wide Association Studies-GWAS
15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.3K
Gene Families
3.5K
3.5K
Gene Families
9.8K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
9.8K
Genome Size and the Evolution of New Genes
3.3K
3.3K
Genome Size and the Evolution of New Genes
9.0K
While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.
9.0K
