文本挖掘生物医学文献以识别数字流行病学和系统审查的极度不平衡数据:SARS-CoV-2基因组流行病学研究数据集和方法
Davy Weissenbacher1, Karen O'Connor2, Ari Klein2
1Cedars-Sinai Medical Center, Los Angeles, CA, USA.
medRxiv : the preprint server for health sciences
|August 14, 2023
概括
使用自然语言处理 (NLP) 的自动化方法识别报告新SARS-CoV-2序列的研究文章. 这有助于从出版物中提取关键的患者数据,以丰富基因组流行病学研究.
科学领域:
- 生物信息学和计算生物学
- 基因组流行病学 基因组流行病学
- 自然语言处理自然语言处理.
背景情况:
- 从科学文献中手动提取数据是耗时和昂贵的,特别是在系统性审查等大规模研究中.
- COVID-19大流行凸显了基因组流行病学研究中需要详细的患者数据的需要,这些数据往往缺少像GenBank和GISAID这样的序列存储库.
- 随着序列数据的发表文章是缺失的高级细节的宝贵来源,例如地理位置和患者人口统计数据.
结论:
- 相关出版物的自动识别是丰富基因组数据库的关键第一步.
- 这种方法大大减少了文献审查和数据提取所需的时间和精力.
- 通过增强数据,实现大规模的基因组流行病学研究,可以揭示病毒变异,传播和患者结果之间的重要关联.
相关概念视频
Genomics
36.5K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.5K
Statistical Software for Data Analysis and Clinical Trials
626
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
626
Single Nucleotide Polymorphisms-SNPs
15.3K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.3K
Statistical Methods for Analyzing Epidemiological Data
411
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
411
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Bias in Epidemiological Studies
352
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
352


