Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Automatic extraction of candidate nomenclature terms using the doublet method.

Jules J Berman1

  • 1Cancer Diagnosis Program, National Cancer Institute, National Institutes of Health, Bethesda, MD, USA. jjberman@alum.mit.edu

BMC Medical Informatics and Decision Making
|October 20, 2005
PubMed
Summary

This study introduces a computational method to identify new biomedical terms for nomenclatures. The doublet coding approach efficiently extracts candidate terms from large text volumes, aiding curators in keeping terminologies current.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

DIFFERENTIATION OF A PRIMARY CHEMICALLY INDUCED RAT NEPHROBLASTOMA IN ORGAN CULTURE.

Development, growth & differentiation·2023
Same author

Post-Informatics pathology.

Journal of pathology informatics·2011
Same author

Informatics research using publicly available pathology data.

Journal of pathology informatics·2011
Same author

The tissue microarray OWL schema: An open-source tool for sharing tissue microarray data.

Journal of pathology informatics·2010
Same author

Minimum information specification for in situ hybridization and immunohistochemistry experiments (MISFISHIE).

Nature biotechnology·2008
Same author

Availability and quality of paraffin blocks identified in pathology archives: a multi-institutional study by the Shared Pathology Informatics Network (SPIN).

BMC cancer·2007

Area of Science:

  • Biomedical Informatics
  • Computational Linguistics
  • Medical Terminology

Background:

  • The biomedical literature continuously generates new terminology, posing challenges for manual curation of existing nomenclatures.
  • Curators require computational tools to efficiently identify and incorporate novel terms amidst the vast volume of published research.
  • Traditional methods of manual literature review are time-consuming and struggle to keep pace with the rapid expansion of biomedical knowledge.

Purpose of the Study:

  • To describe a computational method for the rapid extraction of new, candidate terms from large biomedical text corpora.
  • To assist nomenclature curators by providing an automated system for identifying terms suitable for addition to existing medical terminologies.
  • To demonstrate the efficacy of the doublet coding method for discovering novel nomenclature candidates.

Related Experiment Videos

Main Methods:

  • A variation of the doublet coding method was employed to parse biomedical text, identifying sequences of overlapping word doublets.
  • The algorithm identifies contiguous sequences of word doublets present in a reference nomenclature.
  • Candidate terms are extracted if a matching doublet sequence is not already present in the reference nomenclature.

Main Results:

  • Parsing a 31+ MB corpus of pathology abstracts (4,289 records) yielded 313 candidate new terms in 2 seconds.
  • Human review approved 285 of the 313 candidate terms.
  • A final list of 222 new terms (71% of candidates) was identified for potential addition to the reference nomenclature.

Conclusions:

  • The doublet method provides an efficient and rapid means of extracting candidate nomenclature terms from extensive text datasets.
  • The algorithm is adaptable to virtually any text corpus and nomenclature system.
  • An implementation of the algorithm in Perl is available, facilitating its practical application by researchers and curators.