Related Experiment Video
Updated: Jun 7, 2025

14:06
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
15.2K
Rare copy number variant analysis in case-control studies using snp array data: a scalable and automated data
Haydee Artaza1,2, Ksenia Lavrichenko1,3, Anette S B Wolff1,4
1Department of Clinical Science, University of Bergen, Bergen, Norway.
BMC Bioinformatics
|November 16, 2024
Summary
This study introduces a flexible bioinformatic pipeline for detecting rare copy number variants (CNVs) from SNP array data. The automated framework enhances rare CNV analysis in human genomics research.
Area of Science:
- Genomics
- Bioinformatics
- Human Genetics
Background:
- Rare copy number variants (CNVs) are significant genomic alterations impacting human health and disease susceptibility.
- Analyzing rare CNVs from SNP array data traditionally requires complex bioinformatic pipelines.
- There is a need for efficient and flexible tools to investigate rare CNVs.
Purpose of the Study:
- To develop and present a flexible bioinformatic pipeline for the automated detection and analysis of rare CNVs from human SNP array data.
- To provide a robust framework for researchers investigating the role of rare CNVs in human diseases.
Main Methods:
- The pipeline is implemented using Snakemake, a rule-based workflow management system.
- It comprises two main sub-pipelines: one for variant calling and quality control (QC), and another for rare CNV analysis.
- The framework is designed for automation, scalability, and flexibility.
Main Results:
- The pipeline automates the calling and quality control of CNVs.
- It enables the assessment of rare CNV frequencies in patient versus control cohorts.
- The system facilitates the evaluation of CNV impact on genes and biological pathways.
Conclusions:
- The developed pipeline offers an efficient and flexible bioinformatic solution for rare CNV investigation.
- It incorporates rigorous quality control and comparative frequency analysis for reliable results.
- This framework aims to advance biomedical research by simplifying rare CNV analysis.
Related Concept Videos
Comparing Copy Number Variations and SNPs
17.3K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.3K
Single Nucleotide Polymorphisms-SNPs
14.1K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
14.1K
Genome-wide Association Studies-GWAS
12.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.5K

