Quantifying copy number variations using a hidden Markov model with inhomogeneous emission distributions.
Kenneth Jordan McCallum1, Ji-Ping Wang
1Department of Statistics, Northwestern University, Evanston, IL 60208, USA. kennethmccallum2013@u.northwestern.edu
Biostatistics (Oxford, England)
|February 23, 2013
Summary
Copy number variations (CNVs) detection is improved using a novel hidden Markov model (HMM). This HMM accounts for sequencing biases in high-throughput data, offering competitive performance for disease-associated genetic variation analysis.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Copy number variations (CNVs) are key genetic alterations linked to diseases like cancer and autism.
- High-throughput sequencing generates data for CNV detection, but its statistical properties remain unclear.
- Existing methods for CNV analysis may not fully account for sequencing data biases.
Purpose of the Study:
- To develop a robust statistical model for detecting copy number variations from high-throughput sequencing data.
- To address the distributional properties and sequencing biases inherent in next-generation sequencing data.
- To implement and evaluate a novel algorithm for accurate CNV identification.
Main Methods:
- A hidden Markov model (HMM) was developed with inhomogeneous emission distributions.
- Negative binomial regression was utilized to model sequencing biases within the HMM.
- The proposed model was validated using whole genome sequencing and simulated datasets.
- An R package, CNVfinder, was created to implement the CNV detection algorithm.
Main Results:
- The negative binomial regression-based HMM demonstrated a good fit to sequencing data.
- The CNVfinder algorithm showed competitive performance compared to read count normalization methods.
- The model effectively accounts for biases present in high-throughput sequencing data.
- Accurate detection of copy number variations was achieved.
Conclusions:
- The proposed hidden Markov model provides an effective approach for copy number variation detection.
- Negative binomial regression is a suitable method for modeling biases in sequencing data for CNV analysis.
- The CNVfinder R package offers a competitive and accurate tool for genomic variation studies.
- This work advances the understanding and detection of genetic variations associated with diseases.
Related Concept Videos
Variation: Normal Distribution, Range, and Standard Deviation
27.0K
In the field of psychology, there are several ways to organize measurements of a trait, feature, or characteristic (i.e., variables). Qualitative data, such as ethnicity, can be tabulated into a frequency count to provide information about the proportion, as well as the variety of groups in a sample or population. On the other hand, researchers can perform a wider set of calculations on quantitative data. The mean, mode, and median, for instance, are central tendency measures to identify a...
27.0K
What is Variation?
17.6K
Apart from the measures of central tendency, distribution, outliers, and the changing characteristics of data with time, an important characteristic of any data set is its variation or spread. In some data sets, the data values are concentrated closely near the mean; in others, the data values are more widely spread out from the mean.
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
17.6K
Emission Spectra
75.8K
When solids, liquids, or condensed gases are heated sufficiently, they radiate some of the excess energy as light. Photons produced in this manner have a range of energies, and thereby produce a continuous spectrum in which an unbroken series of wavelengths is present.
75.8K
Variation
7.7K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
7.7K
Conservative Site-specific Recombination and Phase Variation
6.7K
Because the DNA segments are cut and reorganized in a direction-specific manner, site-specific recombination has emerged as an efficient genetic engineering technique. Flippase and Cyclization recombinases or Flp and Cre, respectively, are two members of the tyrosine recombinase family derived from bacteriophages, that are used to mediate site-specific DNA insertions, deletions, and targeted expression of proteins in mammalian cell lines.
The recognition sites for Cre recombinase called LoxP...
The recognition sites for Cre recombinase called LoxP...
6.7K
Quantifying Work
24.1K
As a system undergoes a change, its internal energy can change, and energy can be transferred from the system to the surroundings, or from the surroundings to the system.
24.1K


