Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Combining text mining and sequence analysis to discover protein functional regions.

E Eskin1, E Agichtein

  • 1School of Computer Science Engineering, Hebrew University. eeskin@cs.huji.ac.il

Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
|March 3, 2004
PubMed
Summary

This study introduces a novel method combining text and protein sequence analysis to identify functional protein regions, overcoming data limitations for improved protein classification and localization prediction.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Increasing the power of meta-analysis of genome-wide association studies to detect heterogeneous effects.

Bioinformatics (Oxford, England)·2017
Same author

Genome-wide association study of NMDA receptor coagonists in human cerebrospinal fluid and plasma.

Molecular psychiatry·2015
Same author

Heritability of periodontal bone loss in mice.

Journal of periodontal research·2015
Same author

Genome-wide association study of monoamine metabolite levels in human cerebrospinal fluid.

Molecular psychiatry·2013
Same author

Genome-wide association study of Tourette's syndrome.

Molecular psychiatry·2012
Same author

Genome-wide association study of bipolar disorder in European American and African American individuals.

Molecular psychiatry·2009

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Genomics

Background:

  • Protein sequence classification models can identify functionally relevant regions.
  • Automated identification of these regions is challenging due to data scarcity.
  • Existing methods struggle with limited annotated sequence data.

Purpose of the Study:

  • To address data scarcity in protein sequence analysis by integrating text and sequence data.
  • To develop a robust method for identifying functional regions in protein sequences.
  • To enhance the accuracy of predicting protein sub-cellular localization.

Main Methods:

  • Trained a text classifier on available textual annotations to label unlabeled sequences.
  • Developed a joint sequence-text classifier using extended datasets.

Related Experiment Videos

  • Projected the classifier onto original sequences to pinpoint relevant regions.
  • Main Results:

    • Successfully predicted protein sub-cellular localization.
    • Identified localization-specific functional regions within protein sequences.
    • Demonstrated the effectiveness of the combined text and sequence analysis approach.

    Conclusions:

    • The integrated text and sequence analysis approach effectively overcomes data scarcity challenges.
    • This method enhances the ability to identify functional regions in proteins.
    • The approach shows promise for advancing protein function prediction and classification.