Related Experiment Videos
Evaluation of gene-finding programs on mammalian sequences.
S Rogic1, A K Mackworth, F B Ouellette
1Computer Science Department, The University of California at Santa Cruz, Santa Cruz 95064, USA. rogic@cse.ucsc.edu
Genome Research
|May 5, 2001
Summary
This study evaluates seven gene-finding programs using a novel mammalian genomic dataset. Newer programs show improved accuracy, with analysis detailing strengths and weaknesses for computational gene identification.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Accurate gene identification is crucial for understanding genome function.
- Previous gene-finding program evaluations used datasets that may overlap with program training data.
- A need exists for unbiased evaluation of modern gene-finding algorithms.
Purpose of the Study:
- To independently compare the performance of seven recently developed gene-finding programs.
- To assess gene-finding accuracy based on sequence and prediction features.
- To identify the specific strengths and weaknesses of each program and computational gene-finding overall.
Main Methods:
- Developed a new, biologically validated dataset (HMR195) of mammalian genomic sequences, ensuring no overlap with program training sets.
- Performed an independent comparative analysis of seven gene-finding programs: FGENES, GeneMark.hmm, Genie, Genescan, HMMgene, Morgan, and MZEF.
- Examined prediction accuracy in relation to sequence features (e.g., G+C content) and prediction characteristics (e.g., exon length, signal type).
Main Results:
- The new generation of gene-finding programs demonstrated substantially improved accuracy compared to those in prior studies.
- Analysis revealed varying performance across programs depending on sequence and prediction features.
- Specific strengths and weaknesses of individual programs and the field of computational gene-finding were elucidated.
Conclusions:
- Modern gene-finding programs represent a significant advancement in computational genomics.
- Understanding feature-dependent accuracy is key to selecting appropriate tools for specific genomic analyses.
- The HMR195 dataset and detailed results provide a valuable resource for future gene-finding research and development.