Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Building a protein name dictionary from full text: a machine learning term extraction approach.

Lei Shi1, Fabien Campagne

  • 1Institute for Computational Biomedicine, Dept. of Physiology and Biophysics, Weill Cornell Medical College, 1300 York Ave, New York, NY 10021, USA. les2007@med.cornell.edu

BMC Bioinformatics
|April 9, 2005
PubMed
Summary

This study introduces a new method for identifying biological entity names directly from full-text articles. The approach efficiently creates comprehensive protein name dictionaries, capturing many name variants missed by traditional databases.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

ALK1-BMPRII agonism by clustering bispecific antibodies treats hereditary hemorrhagic telangiectasia.

bioRxiv : the preprint server for biology·2025
Same author

Correcting Smad1/5/8, mTOR, and VEGFR2 treats pathology in hereditary hemorrhagic telangiectasia models.

The Journal of clinical investigation·2019
Same author

Tacrolimus rescues the signaling and gene expression signature of endothelial ALK1 loss-of-function and improves HHT vascular pathology.

Human molecular genetics·2017
Same author

Specific calpain inhibition protects kidney against inflammaging.

Scientific reports·2017
Same author

Glomerular common gamma chain confers B- and T-cell-independent protection against glomerulonephritis.

Kidney international·2017
Same author

A mouse model of hereditary hemorrhagic telangiectasia generated by transmammary-delivered immunoblocking of BMP9 and BMP10.

Scientific reports·2016

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Natural Language Processing

Background:

  • Most biological information is in full-text articles, but data mining tools focus on abstracts.
  • Existing tools use lexicons from curated databases with limited coverage of biological name variants.
  • This gap hinders comprehensive literature data mining.

Purpose of the Study:

  • To develop an efficient and robust approach for recognizing named biological entities directly within full-text scientific articles.
  • To create a more comprehensive dictionary of protein names by extracting terms from full-text literature.
  • To improve the coverage of biological name variants in literature data mining.

Main Methods:

  • Collecting high-frequency terms from full-text articles.

Related Experiment Videos

  • Employing Support Vector Machines (SVM) for the identification of biological entity names.
  • Building a protein name dictionary from a large corpus of 80,528 full-text articles.
  • Main Results:

    • The developed method efficiently recognizes named entities in full text.
    • A protein name dictionary was created, with only 8.3% of names matching SwissProt descriptions.
    • The dictionary demonstrated strong performance in recognizing protein name variants not present in SwissProt.

    Conclusions:

    • The direct extraction approach is significant and compares favorably to existing methods.
    • The SVM-based method is robust and computationally efficient for full-text analysis.
    • This approach enhances the discovery of biological entities and their name variations in scientific literature.