Related Experiment Videos
Automatic discovery of cross-family sequence features associated with protein function
Markus Brameier1, Josien Haan, Andrea Krings
1Stockholm Bioinformatics Center, Stockholm University, 106 91 Stockholm, Sweden. brameier@birc.au.dk
BMC Bioinformatics
|January 18, 2006
Summary
A new self-supervised data mining method discovers protein function from amino acid sequences without predefined categories. It identifies sequence-function links, including novel associations like polyglutamine repeats with transcription.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Predicting protein function from amino acid sequences is crucial for uncharacterized protein families and comparative genomics.
- Current machine learning methods predict predefined functional categories or subcellular locations, potentially missing underlying biological relationships.
- Human-designated functional classes may not perfectly align with biological reality, limiting discovery.
Purpose of the Study:
- To develop a self-supervised data mining approach for discovering protein function directly from sequence data.
- To identify relationships between sequence features and functional annotations without requiring prior biological assumptions.
- To explore novel sequence-to-function correlations beyond established categories.
Main Methods:
- Employed a self-supervised data mining technique utilizing genetic programming.
- Co-evolved amino acid-based regular expressions and keyword-based logical expressions.
- Trained on protein sequences and their associated UniProt/Swiss-Prot annotations.
Main Results:
- Successfully identified relationships between sequence features and functional annotations.
- Detected strong correlations with cellular compartment targeting and broad functional roles like inhibition, biosynthesis, transcription, and defense.
- Found distinct sequence motifs for specific functions, e.g., polyglutamine repeats linked to transcription more than nuclear location.
Conclusions:
- Developed a novel approach for knowledge discovery in annotated sequence data.
- The technique effectively identifies functionally important sequence features without expert knowledge.
- Revealed unexpected links between biological processes, such as ubiquitination and transcription, by analyzing protein function from a sequence perspective.