Text as data: using text-based features for proteins representation and for computational prediction of their

Hagit Shatkay1, Scott Brady2, Andrew Wong3

  • 1Dept. of Computer and Information Sciences, University of Delaware, Newark, DE 19716, USA; Delaware Biotechnology Institute, University of Delaware, Newark, DE 19711, USA; Computational Biology and Machine Learning Lab, School of Computing, Queen's University, Kingston, ON K7L 3N6, Canada.

Summary

This study introduces a novel approach using scientific literature as text-based features for predicting protein characteristics. This method enhances computational protein annotation beyond traditional sequence analysis.