Related Experiment Videos
The probabilities of similarities in DNA sequence comparisons
L D Brooks1, B S Weir, H E Schaffer
1Department of Statistics, North Carolina State University, Raleigh 27695-8203.
Genomics
|October 1, 1988
Summary
Assessing DNA sequence similarity significance is crucial. This study introduces a method using the Queen and Korn algorithm to determine if local DNA similarities are statistically significant, providing a table for analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Statistical Genetics
Background:
- Identifying significant local similarities in DNA sequences is fundamental for understanding gene function and evolutionary relationships.
- Existing methods, including some commercial packages, may not accurately assess the statistical significance of these similarities.
- The Queen and Korn algorithm provides a framework for detecting local sequence alignments.
Purpose of the Study:
- To statistically evaluate the significance of local DNA sequence similarities.
- To illustrate a procedure for assessing significance using the Queen and Korn algorithm.
- To provide a practical tool (a table) for determining the significance of similarity lengths.
Main Methods:
- Defined statistical significance at the 5% level based on the probability of observing a similarity length or greater in random sequences.
- Related the distribution of longest similarity lengths to lengths from specific starting positions.
- Implemented the Queen and Korn algorithm, constructing distributions by combining base blocks for similarity extension.
Main Results:
- Developed a method to calculate critical values for assessing the statistical significance of DNA sequence similarities.
- Generated a table to evaluate the significance of longest similarities in sequences up to 1000 bases.
- Demonstrated that substantial similarity lengths can occur by chance alone.
Conclusions:
- The calculated critical values offer a more reliable assessment of significance compared to expected numbers used in some software.
- The proposed method and accompanying table enhance the accuracy of identifying biologically meaningful DNA similarities.
- This approach aids researchers in distinguishing true biological signals from random occurrences in DNA sequence comparisons.