一个基于随机森林的预测模型来分类BRCA1误解变异:一种用于评估误解突变效应的新方法
Hamed Ka1, Maryam Naghinejad2, Akbar Amirfiroozy2
1Department of Computer Science, Faculty of Mathematics, Statistics, and Computer Science, University of Tabriz, Tabriz, Iran.
Journal of human genetics
|April 18, 2025
概括
使用随机森林的新工具BRCA1-Forest准确地对乳腺和卵巢癌风险的BRCA1基因变异进行了分类. 它提供了解释,在大多数指标中表现优于现有方法,以更好地检测症状前疾病.
科学领域:
- 基因组学和生物信息学
- 癌症遗传学 癌症遗传学
- 机器学习在医学中的应用
背景情况:
- 准确的BRCA1变异分类对于早期发现和预防乳腺癌和卵巢癌至关重要.
- 现有的预测工具往往缺乏解释性,并与特定的变异分类作斗争.
- 需要高性能,可解释的工具来评估BRCA1变异的临床意义.
研究的目的:
- 开发一个准确和可解释的预测工具来分类BRCA1变异的临床意义.
- 改善当前变种分类方法的局限性.
- 加强症状前疾病检测和预防遗传性癌症的预防策略.
主要方法:
- 收集了BRCA1良性和致病性误解变体的数据集.
- 通过分析蛋白质序列的变异效应,包括物理化学变化和保存分数来准备数据集.
- 训练了一种基于森林的机器学习模型,BRCA1-Forest,用于变种分类.
主要成果:
- 与SIFT,PolyPhen2,CADD和DANN相比,BRCA1-Forest在多个评估指标 (精度,FPR,AUC ROC,AUC-PR,MCC) 上表现优越.
- 该模型在分类BRCA1变异显著性方面实现了高特异性和灵敏性.
- 在所有测试指标中表现优于现有方法,除了回忆.
结论:
- 在准确和可解释的BRCA1变体分类方面,BRCA1-Forest提供了显著的进步.
- 该工具有可能改善患有遗传性乳腺癌和卵巢癌风险的患者的临床决策.
- 开发的软件可供公众使用,用于研究和临床应用.
相关概念视频
Comparing Copy Number Variations and SNPs
19.4K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
19.4K
Single Nucleotide Polymorphisms-SNPs
20.1K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
20.1K
Expected Frequencies in Goodness-of-Fit Tests
8.9K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
8.9K
Prediction Intervals
3.6K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.6K
Principles of Pharmacogenetics: Types of Genetic Variants
113
The human genome is over 99.9% identical between individuals, yet genetic differences exist at millions of bases. The human genome contains approximately 3 million variant positions per individual, many of which are heterozygous, contributing to genetic diversity and individual traits. Genetic variations include single-nucleotide polymorphisms (SNPs), insertions, deletions, and copy number variations (CNVs).SNPs, the most common variation, involve single-base changes in DNA. These can be...
113
Cancer Survival Analysis
846
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
846


