Constructing germline research cohorts from the discarded reads of clinical tumor sequences

Alexander Gusev1,2,3, Stefan Groha4,5, Kodi Taraszka6

  • 1Division of Population Sciences, Dana-Farber Cancer Institute and Harvard Medical School, Boston, MA, USA. alexander_gusev@dfci.harvard.edu.

Genome Medicine
|November 9, 2021
PubMed
Abstract

Insights

Discarded tumor sequencing data can reveal genome-wide germline genotypes, enabling new research into cancer genetics and patient risk factors. This approach enhances germline genetic insights from existing tumor sequencing archives.

Area of Science:

  • Genomics
  • Cancer Research
  • Bioinformatics

Background:

  • Targeted tumor sequencing is widely used in oncology to identify mutations and has yielded significant discoveries.
  • However, its targeted nature limits its utility for comprehensive germline genetic analysis.
  • Off-target reads from tumor sequencing are typically discarded, representing a missed opportunity for germline genetic research.

Purpose of the Study:

  • To assess the utility of discarded off-target reads from tumor-only panel sequencing for recovering genome-wide germline genotypes.
  • To develop and validate a computational framework for inferring germline genetic information from tumor sequencing data.
  • To demonstrate the feasibility and value of leveraging existing tumor sequencing data for germline genetic research.

Main Methods:

  • Developed a computational framework for germline variant inference, including imputation, quality control, genetic ancestry estimation, polygenic risk score calculation, and HLA allele inference.
  • Benchmarked the framework using tumor sequencing data from 833 individuals with matched germline SNP array data.
  • Applied the validated framework to a large cohort of 25,889 prospectively collected tumor samples.

Main Results:

  • Demonstrated high to moderate accuracy for imputed common variants (correlation 0.86), inferred genetic ancestry (>0.98), polygenic risk scores (>0.90), and HLA alleles (>0.80) compared to direct genotyping.
  • Showcased minimal impact on the accuracy of somatic copy number alterations and other tumor features.
  • Identified novel relationships between genetic ancestry, polygenic risk, and tumor characteristics in the 25,889-tumor cohort.

Conclusions:

  • Targeted tumor sequencing data can be effectively repurposed to create extensive germline research cohorts.
  • The developed analysis pipeline is publicly available to facilitate broader adoption and research.
  • This approach significantly expands the utility of existing clinical tumor sequencing data for fundamental and translational research.

Related Concept Videos