Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Protein names precisely peeled off free text.

Sven Mika1, Burkhard Rost

  • 1CUBIC, Department of Biochemistry and Molecular Biophysics, Columbia University, New York, NY 10032, USA. mika@cubic.bioc.columbia.edu

Bioinformatics (Oxford, England)
|July 21, 2004
PubMed
Summary

A new system, NLProt, accurately identifies protein names in scientific literature using support vector machines (SVMs). This method improves upon existing approaches for data mining and information extraction from biomedical texts.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

On the state of protein function prediction: a report on the fourth CAFA challenge.

bioRxiv : the preprint server for biology·2026
Same author

Advances in Protein Function Prediction from the Fifth CAFA Challenge.

bioRxiv : the preprint server for biology·2026
Same author

Whole-genome prediction of bacterial pathogenic capacity on novel bacteria using protein language models with PathogenFinder2.

Bioinformatics (Oxford, England)·2026
Same author

Biocentral: Embedding-based Protein Predictions.

Journal of molecular biology·2026
Same author

Toxin data quality: a critical examination of bacterial exotoxins and animal toxins.

BMC research notes·2025
Same author

FlatProt: 2D visualization eases protein structure comparison.

BMC bioinformatics·2025

Area of Science:

  • Biomedical Informatics
  • Computational Biology
  • Natural Language Processing

Background:

  • Automated protein name identification is crucial for data mining scientific literature.
  • Current methods rely on dictionaries, rules, and machine learning.

Purpose of the Study:

  • To introduce NLProt, a novel system for identifying protein names in MEDLINE abstracts.
  • To combine dictionary/rule-based filtering with support vector machines (SVMs).

Main Methods:

  • A hybrid approach using pre-processing filters and multiple trained SVMs.
  • Training on a corpus of 200 annotated abstracts.
  • Developing guidelines to address redundancy in evaluation sets.

Main Results:

Related Experiment Videos

  • NLProt achieved 75% precision and 76% recall in protein name extraction.
  • The system demonstrated superior performance compared to other methods when evaluated with new guidelines.
  • SVMs effectively utilized the local context of protein names.

Conclusions:

  • NLProt represents a precise and general method for protein name tagging.
  • The system enhances the ability to data mine the wealth of information in scientific literature.
  • Improved evaluation strategies are necessary for reliable performance assessment.