Related Experiment Videos
Transcriptome analyses of human genes and applications for proteome analyses
1Laboratory of Functional Genomics, Department of Medical Genome Sciences, Graduate School of Frontier Sciences, The University of Tokyo, 4-6-1 Shirokanedai, Minatoku, Tokyo 108-8639, Japan. ysuzuki@ims.u-tokyo.ac.jp
Current Protein & Peptide Science
|April 14, 2006
Summary
Full-length cDNA sequencing projects have identified novel human gene transcripts and alternative splicing. Expert annotation, like the H-invitational meeting, is crucial for accurate protein coding region determination and understanding gene networks.
Area of Science:
- Genomics
- Molecular Biology
- Bioinformatics
Background:
- Recent advances in full-length cDNA technologies have enabled large-scale sequencing of human gene transcripts.
- Existing cDNA resources now encompass a significant portion of protein-coding human genes, revealing novel transcripts and splice isoforms.
Purpose of the Study:
- To address challenges in accurately identifying protein-coding regions and translation initiation sites from full-length cDNA sequences.
- To establish a comprehensive functional annotation of full-length cDNAs through international collaboration.
- To leverage transcriptome data for a deeper understanding of human gene networks at the proteome level.
Main Methods:
- Large-scale full-length cDNA sequencing and data collection from global projects.
- Comprehensive bioinformatic analysis of cDNA sequences to identify novel transcripts and isoforms.
- International expert-driven manual and computational annotation of cDNAs (H-invitational meeting).
- Proteome-wide mass-spectrometry analysis to identify small proteins and upstream open reading frames.
Main Results:
- Identification of thousands of novel human gene transcripts and alternatively spliced isoforms.
- Demonstration that the longest open reading frame or the first ATG codon do not always represent the true protein-coding region.
- Discovery of a significant population of small proteins encoded by upstream open reading frames.
- Creation of an integrated database of functional annotations for full-length cDNAs.
- Validation of full-length cDNA data utility in identifying alternative splicing, transcription start sites, and promoter regions.
Conclusions:
- Accurate functional annotation of full-length cDNAs requires expert manual curation alongside computational methods.
- Full-length cDNA data is a valuable resource for understanding gene structure, regulation, and alternative splicing.
- Integrated transcriptome information, combined with genome data, is foundational for achieving a proteome-level understanding of human gene networks.