Related Experiment Video
Updated: Dec 31, 2025

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
New Approximate Statistical Significance of Gapped Alignments Based on the Greedy Extension Model
Amirhossein Karami1, Afshin Fayyaz Movaghar1, Sabine Mercier2
1Department of Statistics, Faculty of Mathematical Sciences, University of Mazandaran, Babolsar, Iran.
Abstract:
Sequence alignment is a fundamental concept in bioinformatics to distinguish regions of similarity among various sequences. The degree of similarity has been considered as a score. There are a number of various methods to find the statistical significance of similarity in the gapped and ungapped cases. In this article, we improve the statistical significance accuracy of the local score by introducing a new approximate p-value. This is developed according to Poisson clumping and the exact distribution of a partial sum of random variables. The efficiency of the proposed method is compared with that of previous methods on real and simulated data. The results yield a remarkable improvement in accuracy of the p-value in the gapped case. This is an evidence for the method to be considered as a prospective candidate for sequences comparison.
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
Wald-Wolfowitz Runs Test II
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
Evolutionary Relationships through Genome Comparisons
Gene Duplication and Divergence
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
Statistical Significance
Expected Frequencies in Goodness-of-Fit Tests

