Related Experiment Video
Updated: Feb 4, 2026

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
A weighted burden test using logistic regression for integrated analysis of sequence variants, copy number variants
David Curtis1,2
1Centre for Psychiatry, Barts and the London School of Medicine and Dentistry, Charterhouse Square, London, EC1M 6BQ, UK. d.curtis@ucl.ac.uk.
Abstract:
Previously described methods of analysis allow variants in a gene to be weighted more highly according to rarity and/or predicted function and then for the variant contributions to be summed into a gene-wise risk score, which can be compared between cases and controls using a t-test. However, this does not allow incorporating covariates into the analysis. Schizophrenia is an example of an illness where there is evidence that different kinds of genetic variation can contribute to risk, including common variants contributing to a polygenic risk score (PRS), very rare copy number variants (CNVs) and sequence variants. A logistic regression approach has been implemented to compare the gene-wise risk scores between cases and controls, while incorporating as covariates population principal components, the PRS and the presence of pathogenic CNVs and sequence variants. A likelihood ratio test is performed, comparing the likelihoods of logistic regression models with and without this score. The method was applied to an ethnically heterogeneous exome-sequenced sample of 6000 controls and 5000 schizophrenia cases. In the raw analysis, the test statistic is inflated but inclusion of principal components satisfactorily controls for this. In this dataset, the inclusion of the PRS and effect from CNVs and sequence variants had only small effects. The set of genes which are FMRP targets showed some evidence for enrichment of rare, functional variants among cases (p = 0.0005). This approach can be applied to any disease in which different kinds of genetic and non-genetic risk factors make contributions to risk.
Related Concept Videos
Histone Variants at the Centromere
Polygenic Traits
Regression Toward the Mean
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Microsoft Excel: Regression Analysis
To perform regression...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...

