Related Experiment Videos
Computational analysis of full-length mouse cDNAs compared with human genome sequences
S Kondo1, A Shinagawa, T Saito
1Laboratory for Genome Exploration Research Group, RIKEN Genomic Sciences Center (GSC), RIKEN Yokohama Institute, 1-7-22 Suehiro-cho, Tsurumi-ku, Yokohama City, Kanagawa 230-0045, Japan. rgscerg@gsc.riken.go.jp
Summary
Identifying novel human genes is challenging. This study uses mouse complementary DNAs (cDNAs) to identify new human gene candidates, revealing potential novel genes and improving gene structure analysis.
Area of Science:
- Genomics
- Molecular Biology
- Bioinformatics
Background:
- Human genome sequencing is complete, but gene identification and structure determination remain challenging.
- Computational gene prediction methods like GENSCAN have limitations in identifying certain exon types.
Purpose of the Study:
- To develop and apply a method using full-length mouse complementary DNAs (cDNAs) to aid in human gene identification and structure determination.
- To identify potential novel human genes and analyze genomic regions.
- To evaluate the performance of gene prediction tools.
Main Methods:
- Alignment of 61,227 RIKEN mouse cDNAs (full-length and ESTs) against draft human genome sequences.
- Identification of 35,141 non-redundant genomic regions with significant mouse cDNA alignment.
- Analysis of genomic regions using full-length cDNAs, including cross-species comparisons.
- Evaluation of GENSCAN prediction biases for exon size and GC-content.
Main Results:
- 35,141 significant alignments between mouse cDNAs and human genomic regions were identified.
- Analysis revealed GENSCAN's systematic bias against small or low GC-content exons.
- 3,217 cDNAs mapped to these regions did not match known human genes or ESTs.
- 1,141 of these cDNAs showed no significant protein similarity, indicating potential novel genes.
Conclusions:
- Full-length mouse cDNAs are effective tools for human gene identification and structure analysis.
- The study identified over a thousand candidate novel human genes.
- Findings highlight limitations in current gene prediction algorithms and suggest areas for improvement.