Related Experiment Videos
Gene sequence signatures revealed by mining the UniGene affiliation network
Jiexin Zhang1, Li Zhang, Kevin R Coombes
1Department of Biostatistics and Applied Mathematics, The University of Texas M.D. Anderson Cancer Center, 1515 Holcombe Boulevard, Box 447, Houston, TX 77030-4009, USA.
Bioinformatics (Oxford, England)
|December 13, 2005
Summary
Most narrowly expressed transcripts (NETs) in UniGene clusters resemble intergenic DNA sequences, questioning their quality. However, some NETs share features with prevalently expressed transcripts (PETs), suggesting fewer quality issues and aiding non-coding RNA studies.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- The post-genomic era necessitates tools for interpreting genomic sequence data.
- UniGene clusters (UCs) classify gene transcripts based on expression patterns.
- This study focuses on narrowly expressed transcripts (NETs) and prevalently expressed transcripts (PETs).
Purpose of the Study:
- To investigate and compare gene sequences of NETs and PETs using affiliation network theory.
- To identify distinguishing features between NETs and PETs.
- To develop a model for differentiating transcript types.
Main Methods:
- Exploration of human and mouse UniGene databases.
- Comparative analysis of NETs and PETs based on sequence characteristics and expression levels.
- Dinucleotide frequency analysis to differentiate transcript types.
- Development of a discriminant analysis model.
Main Results:
- NETs exhibit smaller cluster size, shorter sequences, fewer LocusLink annotations, and lower, sporadic expression compared to PETs.
- Dinucleotide frequencies of NETs are similar to intergenic sequences, differing from PETs.
- A discriminant analysis model was developed based on dinucleotide frequencies to distinguish PETs from intergenic sequences.
Conclusions:
- Most NETs resemble intergenic sequences, raising concerns about UniGene cluster quality.
- A subset of NETs shares features with PETs, indicating potentially higher quality.
- Findings support non-coding RNA research and gene sequence database validation.