在肺癌中进行的正式统计复制分析,全基因组协会研究
Yung-Han Chang1, Jinyoung Byun2,3,4, Bryan R Gorman5
1Department of Biostatistics, University of Texas MD Anderson Cancer Center UTHealth Houston Graduate School of Biomedical Sciences, Houston, TX, USA.
medRxiv : the preprint server for health sciences
|November 19, 2025
概括
基于统计模型的复制分析显著减少了肺癌全基因组关联研究 (GWAS) 的错误阳性. 这种方法更有效地识别了关键单核酸多态 (SNP),改善了多基因风险评分的开发.
科学领域:
- 遗传学 遗传学 是一个
- 癌症研究 癌症研究
- 统计基因组学 统计基因组学
背景情况:
- 全基因组关联研究 (GWAS) 已经确定了许多与肺癌风险相关的单核酸多态 (SNP).
- 将GWAS发现转化为临床应用受到高假阳性率 (I型错误) 的阻碍.
- 现有的方法,如p值值和元分析,提供有限的虚假关联的减少.
研究的目的:
- 引入和验证基于统计模型的复制分析,用于从GWAS中策划高质量的显著SNP.
- 为了比较基于模型的复制与传统的元分析在减少错误发现方面的有效性.
- 评估基于复制的SNP衍生出的多基因风险评分 (PRS) 的性能.
主要方法:
- 开发了复制复合式零假设的正式统计测试,确保跨队列一致的SNP效应方向.
- 进行了双向和三向模拟,以评估与元分析相比的错误发现率 (FDR).
- 来自国际肺癌联盟GWAS的复制SNP用于状细胞肺癌和肺腺癌.
- 使用基于复制的和GWAS显著的SNP构建的多基因风险得分 (PRS).
主要成果:
- 基于模型的复制分析表明,错误发现率 (FDR) 比元分析低得多 (在双向模拟中低6.4倍).
- 在三向复制中,9.8%的GWAS显著SNP被验证为状细胞肺癌和33.8%的肺腺癌.
- 基于复制的PRS的性能与GWAS显著的PRS相提并论,但使用的变体少87.3%.
结论:
- 正式的基于模型的复制分析有效地减少了来自GWAS的虚假发现.
- 这种方法提高了将遗传发现转化为临床见解的稳定性和效率.
- 基于模型的复制有助于开发更有效,更准确的多基因肺癌风险评分.
相关概念视频
Genome-wide Association Studies-GWAS
15.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.2K
Statistical Methods for Analyzing Epidemiological Data
885
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
885
Comparing Copy Number Variations and SNPs
18.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.5K


