GPAD:基于自然语言处理的应用程序,从OMIMIM中提取基因疾病关联发现信息
K M Tahsin Hassan Rahit1,2, Vladimir Avramovic1,2, Jessica X Chong3,4
1Departments of Biochemistry, Molecular Biology and Medical Genetics, Cumming School of Medicine, University of Calgary, Calgary, AB, T2N 4N1, Canada.
BMC bioinformatics
|February 27, 2024
概括
一个新的工具,基因表型协会发现 (GPAD),使用自然语言处理从OMIM中提取基因疾病协会. 尽管发现后外体序列的发现增加,但近年来显示下降,需要更大的队列和增加模型生物的使用.
科学领域:
- 遗传学和生物信息学
- 计算生物学 计算生物学
背景情况:
- 成千上万的基因与孟德尔条件有关,在线孟德尔人类遗传 (OMIM) 数据库是关键资源.
- OMIM数据主要是文本和异质的,这给自动数据提取和分析带来了挑战.
- 自然语言处理 (NLP) 提供了一种处理和从复杂的生物文本中提取结构信息的解决方案.
研究的目的:
- 开发一种工具,基因-表型协会发现 (GPAD),用于从OMIM中自动提取基因-疾病协会 (GDA).
- 分析GDA发现的趋势,并确定影响发现率的因素.
- 为研究战略规划提供GDA发现的实时分析.
主要方法:
- 利用自然语言处理 (NLP) 和基于语言的技术来处理OMIM API中的文本.
- 开发了基因表型协会发现 (GPAD) 工具以提取GDA信息,包括验证方法 (模型生物,患者队列).
- 将GPAD提取的数据与已发表的报告进行验证,并将其与大型语言模型性能进行比较.
主要成果:
- GPAD成功地提取了GDA信息,详细介绍了基因表型链接和验证类型.
- 观察到,在引入外基因组测序后,GDA发现率显著增加.
- 确定了2017-2022年GDA发现率的下降,这与需要更大的队列和增加使用斑马鱼和Drosophila等模型生物的需求有关.
结论:
- GPAD提供了GDA发现的实时分析,有助于研究战略和管理.
- 该工具的容量可以扩展到从OMIM和科学文献中获取额外的信息.
- 这些发现突出了基因研究中的不断变化的趋势,以及计算工具在数据分析中的重要性.
更多相关视频
09:37Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
Published on: August 15, 2019
9.7K
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
33.8K
相关概念视频
Genomics
36.3K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.3K
Genome-wide Association Studies-GWAS
13.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.4K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Single Nucleotide Polymorphisms-SNPs
15.1K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.1K
