Related Experiment Video
Updated: Feb 25, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Unsupervised Medical Subject Heading Assignment Using Output Label Co-occurrence Statistics and Semantic Predications
Ramakanth Kavuluru1,2, Zhenghao He2
1Division of Biomedical Informatics, Department of Biostatistics, University of Kentucky, Lexington, KY.
This study introduces an unsupervised method for indexing biomedical articles using Medical Subject Headings (MeSH) terms. It leverages text analysis and term co-occurrence, achieving results comparable to supervised machine learning approaches.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Medical Subject Headings (MeSH)
Background:
- The National Library of Medicine (NLM) uses Medical Subject Headings (MeSH) for indexing biomedical literature.
- Automated methods for MeSH term assignment often rely on supervised machine learning, requiring extensive labeled training data.
- Biomedical data privacy concerns limit the availability of such training datasets.
Purpose of the Study:
- To develop and evaluate a novel unsupervised approach for assigning MeSH terms to biomedical abstracts.
- To explore the utility of named entity recognition, relationship extraction, and MeSH term co-occurrence frequencies.
- To assess the potential of unsupervised methods in overcoming data limitations in biomedical indexing.
Main Methods:
- Utilized named entity recognition and relationship extraction on a large corpus of NLM-indexed articles.
- Calculated co-occurrence frequencies for MeSH term pairs from existing indexed articles.
- Developed an unsupervised model based on extracted relationships and co-occurrence statistics.
Main Results:
- The unsupervised approach achieved a micro F-score comparable to supervised methods.
- Demonstrated the effectiveness of using output label co-occurrence statistics for indexing.
- Showcased the feasibility of unsupervised term recommendation without sensitive training data.
Conclusions:
- Unsupervised methods utilizing text-derived relationships and label co-occurrences show significant promise for biomedical article indexing.
- This approach offers a viable alternative to supervised methods, especially when labeled data is scarce or sensitive.
- Further research into exploiting free-text relationships and co-occurrences can enhance automated indexing systems.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Anatomical Terminology
Functional Classification of Joints
The functional classification of joints is determined by the amount of mobility between the adjacent bones. Joints are functionally classified as a synarthrosis or immobile joint, an amphiarthrosis or slightly moveable joint, or as a diarthrosis, a freely moveable joint. Fibrous and cartilaginous joints can be functionally classified as either synarthroses or amphiarthroses, whereas all synovial joints are classified as diarthroses.
Synarthrosis
An...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Classification of Leukocytes
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
