Related Experiment Videos
On the dependence structure of sequence alignment scores calculated with multiple scoring matrices.
Florian Frommlet1, Andreas Futschik
1Medical University of Vienna.
Summary
Using multiple scoring matrices for protein sequence alignment creates interpretation issues for p-values. This study introduces logistic copula functions to accurately correct p-values and measure score dependence between matrices.
Area of Science:
- Bioinformatics
- Computational Biology
- Statistical Modeling
Background:
- Protein sequence alignment is crucial for understanding protein function and evolution.
- Current methods often involve testing multiple scoring matrices, leading to statistical challenges.
- Interpreting p-values and E-values becomes difficult due to the multiple testing problem.
Purpose of the Study:
- To address the multiple testing problem in protein sequence alignment when using various scoring matrices.
- To develop a statistically sound method for interpreting alignment scores derived from different matrices.
- To introduce logistic copula functions for modeling score dependence.
Main Methods:
- Focusing on local alignment algorithms.
- Employing logistic copula functions to explicitly model the dependence structure between scores from different matrices.
- Deriving p-value correction factors for the use of multiple scoring matrices.
Main Results:
- The proposed method provides accurate p-value correction factors when multiple scoring matrices are applied to the same sequences.
- The logistic copula parameter quantifies the dependence between scores from different matrices.
- This offers insights into the relatedness of scoring matrices.
Conclusions:
- Logistic copula functions offer a robust statistical framework for protein sequence alignment.
- Accurate p-value interpretation is achievable even when employing multiple scoring matrices.
- The dependence measure provides valuable information about scoring matrix similarity.