使用种群遗传学方法对变种病原性预测方法的基准测试
Mikhail Gudkov1,2, Loïc Thibaut3, Steven Monger2
1Victor Chang Cardiac Research Institute, Darlinghurst, NSW 2010, Australia.
Bioinformatics advances
|November 4, 2025
概括
变异性致病性预测因子对于罕见疾病研究至关重要. 我们的研究确定了CADD和REVEL作为使用人口数据的顶级预测者,避免了传统方法的偏见.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 变异性致病性预测因子对于罕见疾病研究至关重要.
- 目前的预测者面临的挑战是由于在培训/测试数据集中的确定偏差和数据循环性.
- 存在需要可靠的基准测试方法,独立于预定义的"基础真相"变量集.
研究的目的:
- 使用直角方法对常用的变异性致病性预测因子进行基准测试.
- 确定最可靠的预测因素,以区分有害的遗传变异.
- 评估"情境调整的单身比例 (CAPS) "指标的有用性.
主要方法:
- 使用gnomAD.使用人口层次的基因组数据对病原性预测者的基准测试.
- 在变异分析中使用背景调整的单子比例 (CAPS) 度量.
- 采用了正交的方法,避免依赖于精选的疾病或突变性变异组.
主要成果:
- 确定CADD和REVEL是区分变体有害性的表现最好的预测指标.
- 在评估的预测指标中,REVEL显示出优异的校准.
- CAPS证明了作为元分析工具的实用性,并突出了基于ClinVar的培训中的偏见.
结论:
- 这项研究提供了一种强大的方法来评估变异性病原性预测因子.
- 建议CADD和REVEL用于识别有害变异,REVEL提供更好的校准.
- 在预测器培训中,CAPS为基于人口的变体解释和偏差检测提供了一个有价值的工具.
相关概念视频
Evolutionary Relationships through Genome Comparisons
6.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.8K
Modern Molecular Taxonomy
574
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
574
Comparing Copy Number Variations and SNPs
18.6K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.6K
Gene Flow
37.4K
Gene flow is the transfer of genes among populations, resulting from either the dispersal of gametes or from the migration of individuals.
37.4K
What is Population Genetics?
64.3K
A population is composed of members of the same species that simultaneously live and interact in the same area. When individuals in a population breed, they pass down their genes to their offspring. Many of these genes are polymorphic, meaning that they occur in multiple variants. Such variations of a gene are referred to as alleles. The collective set of all the alleles within a population is known as the gene pool.
64.3K
Single Nucleotide Polymorphisms-SNPs
17.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.9K


