STREAM-PRS:一个多工具管道,用于简化多基因风险评分计算
Sara Becelaere1,2, Yasmina Abakkouy1, Deborah Sarah Jans1
1Laboratory for Complex Genetics, Department of Human Genetics, KU Leuven, Leuven, Belgium.
Genome medicine
|October 10, 2025
概括
通过测试多种工具,STREAM-PRS优化了多基因风险评分 (PRS) 计算,确定了疾病预测的最佳策略. 这条管道提高了PRS可移植性和基因倾向评估的准确性.
科学领域:
- 遗传学和基因组学 遗传学和基因组学
- 计算生物学 计算生物学
- 疾病风险预测 疾病风险预测
背景情况:
- 多基因风险评分 (PRS) 估计了对疾病的遗传倾向,但存在许多具有不同策略的计算工具.
- 人口分层和PRS可移植性等挑战需要强大的工具和设置选择方法.
- STREAM-PRS是作为一个管道来系统地评估和选择最佳的PRS计算工具而开发的.
研究的目的:
- 开发和验证STREAM-PRS,这是一个有效的管道,用于选择最佳多基因风险评分 (PRS) 计算工具和设置.
- 为应对跨不同数据集和研究中心的PRS可移植性的挑战.
- 应用管道来确定炎症性肠病 (IBD) 的最佳PRS.
主要方法:
- STREAM-PRS使用五种流行的工具 (PRSice-2,PRS-CS,LDpred2,lassosum,lassosum2) 计算PRS,这些工具具有各种设置.
- 管道在培训数据集中选择最佳变体,并在测试数据集中计算分数,然后进行PC校正和标准化.
- 性能是基于差异解释 (R2) 进行评估的;适用于内部IBD队列,使用1000 Genomes数据进行培训和英国生物库数据进行测试.
主要成果:
- 在大约20小时内,STREAM-PRS在5个工具中完成了472个PRS计算.
- 对于IBD,具有特定收缩 (0.7) 的拉索索和兰巴达 (0.008859) 值表现最好,在验证中实现了0.203的R2和0.75的AUC.
- 优化的PRS有效地识别了高风险个体 (PPV=0.905),但在排除低风险个体 (NPV=0.341) 方面不那么可靠.
结论:
- STREAM-PRS为优化PRS计算策略和改善跨中心可移植性提供了一个高效的框架.
- 管道有助于为特定的研究问题选择最合适的PRS工具和设置.
- STREAM-PRS可以在 https://github.com/SaraBecelaere/STREAM-PRS.上公开使用.
相关概念视频
Polygenic Traits
68.9K
When more than one gene is responsible for a given phenotype, the trait is considered polygenic. Human height is a polygenic trait. Studies have uncovered hundreds of loci that influence height, and there are believed to be many more. Due to the high number of genes involved, as well as environmental and nutritional factors, height varies significantly within a given population. The distribution of height forms a bell-shaped curve, with relatively few individuals in the population at the...
68.9K
Single Nucleotide Polymorphisms-SNPs
17.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.9K
Genome-wide Association Studies-GWAS
15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.3K
Multiple Allele Traits
37.9K
The Concept of Multiple Allelism
37.9K
Comparing Copy Number Variations and SNPs
18.6K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.6K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K


