Related Experiment Video
Updated: Apr 20, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Text as data: using text-based features for proteins representation and for computational prediction of their
Hagit Shatkay1, Scott Brady2, Andrew Wong3
1Dept. of Computer and Information Sciences, University of Delaware, Newark, DE 19716, USA; Delaware Biotechnology Institute, University of Delaware, Newark, DE 19711, USA; Computational Biology and Machine Learning Lab, School of Computing, Queen's University, Kingston, ON K7L 3N6, Canada.
This study introduces a novel approach using scientific literature as text-based features for predicting protein characteristics. This method enhances computational protein annotation beyond traditional sequence analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- The rapid expansion of genomic data necessitates advanced methods for protein function and localization prediction.
- Scientific literature is a vast, underutilized resource for protein information, growing exponentially each year.
- Current computational tools often rely solely on protein sequence data for annotation.
Purpose of the Study:
- To explore the utility of scientific literature as a source of text-based features for protein characterization.
- To develop and evaluate a novel computational approach for predicting protein subcellular location and function.
- To demonstrate the value of literature-derived features in complementing sequence-based methods.
Main Methods:
- Developing computational tools to extract text-based features from scientific publications.
- Integrating these text-based features with traditional sequence-based features.
- Applying machine learning models to predict protein subcellular location and function using combined feature sets.
Main Results:
- Text-based features derived from literature significantly improve the accuracy of protein subcellular location prediction.
- The approach shows promise for predicting protein function, offering a complementary perspective to sequence analysis.
- Demonstrated the effectiveness of using literature as a feature source in large-scale biological data analysis.
Conclusions:
- Scientific literature represents a valuable, untapped resource for computational protein annotation.
- This text-mining-based approach offers a powerful complement to existing sequence-based methods.
- The methodology has broad potential for advancing our understanding of protein properties in the era of big data.
More Related Videos
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Related Concept Videos
Protein Organization
The primary structure of a protein is its amino acid sequence....
Protein Organization
Protein Organization
Protein Organization
Protein and Protein Structures
Conservation of Protein Domains