Related Experiment Video
Updated: Oct 14, 2025

Comparative Lesions Analysis Through a Targeted Sequencing Approach
Published on: November 5, 2019
Constructing germline research cohorts from the discarded reads of clinical tumor sequences
Alexander Gusev1,2,3, Stefan Groha4,5, Kodi Taraszka6
1Division of Population Sciences, Dana-Farber Cancer Institute and Harvard Medical School, Boston, MA, USA. alexander_gusev@dfci.harvard.edu.
Background:
Hundreds of thousands of cancer patients have had targeted (panel) tumor sequencing to identify clinically meaningful mutations. In addition to improving patient outcomes, this activity has led to significant discoveries in basic and translational domains. However, the targeted nature of clinical tumor sequencing has a limited scope, especially for germline genetics. In this work, we assess the utility of discarded, off-target reads from tumor-only panel sequencing for the recovery of genome-wide germline genotypes through imputation.
Methods:
We developed a framework for inference of germline variants from tumor panel sequencing, including imputation, quality control, inference of genetic ancestry, germline polygenic risk scores, and HLA alleles. We benchmarked our framework on 833 individuals with tumor sequencing and matched germline SNP array data. We then applied our approach to a prospectively collected panel sequencing cohort of 25,889 tumors.
Results:
We demonstrate high to moderate accuracy of each inferred feature relative to direct germline SNP array genotyping: individual common variants were imputed with a mean accuracy (correlation) of 0.86, genetic ancestry was inferred with a correlation of > 0.98, polygenic risk scores were inferred with a correlation of > 0.90, and individual HLA alleles were inferred with a correlation of > 0.80. We demonstrate a minimal influence on the accuracy of somatic copy number alterations and other tumor features. We showcase the feasibility and utility of our framework by analyzing 25,889 tumors and identifying the relationships between genetic ancestry, polygenic risk, and tumor characteristics that could not be studied with conventional on-target tumor data.
Conclusions:
We conclude that targeted tumor sequencing can be leveraged to build rich germline research cohorts from existing data and make our analysis pipeline publicly available to facilitate this effort.
Insights
Discarded tumor sequencing data can reveal genome-wide germline genotypes, enabling new research into cancer genetics and patient risk factors. This approach enhances germline genetic insights from existing tumor sequencing archives.
Area of Science:
- Genomics
- Cancer Research
- Bioinformatics
Background:
- Targeted tumor sequencing is widely used in oncology to identify mutations and has yielded significant discoveries.
- However, its targeted nature limits its utility for comprehensive germline genetic analysis.
- Off-target reads from tumor sequencing are typically discarded, representing a missed opportunity for germline genetic research.
Purpose of the Study:
- To assess the utility of discarded off-target reads from tumor-only panel sequencing for recovering genome-wide germline genotypes.
- To develop and validate a computational framework for inferring germline genetic information from tumor sequencing data.
- To demonstrate the feasibility and value of leveraging existing tumor sequencing data for germline genetic research.
Main Methods:
- Developed a computational framework for germline variant inference, including imputation, quality control, genetic ancestry estimation, polygenic risk score calculation, and HLA allele inference.
- Benchmarked the framework using tumor sequencing data from 833 individuals with matched germline SNP array data.
- Applied the validated framework to a large cohort of 25,889 prospectively collected tumor samples.
Main Results:
- Demonstrated high to moderate accuracy for imputed common variants (correlation 0.86), inferred genetic ancestry (>0.98), polygenic risk scores (>0.90), and HLA alleles (>0.80) compared to direct genotyping.
- Showcased minimal impact on the accuracy of somatic copy number alterations and other tumor features.
- Identified novel relationships between genetic ancestry, polygenic risk, and tumor characteristics in the 25,889-tumor cohort.
Conclusions:
- Targeted tumor sequencing data can be effectively repurposed to create extensive germline research cohorts.
- The developed analysis pipeline is publicly available to facilitate broader adoption and research.
- This approach significantly expands the utility of existing clinical tumor sequencing data for fundamental and translational research.

