Related Experiment Video
Updated: Jul 15, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Multilabel associative classification categorization of MEDLINE articles into MeSH keywords
Rafal Rak1, Lukasz A Kurgan, Marek Reformat
1Department of Electrical and Computer Engineering, University of Alberta, Edmonton, Canada. rrak@ece.ualberta.ca
A new data mining system effectively classifies MEDLINE documents into multiple Medical Subject Headings (MeSH) categories. This novel approach achieves 90% accuracy, outperforming previous methods for large-scale medical text analysis.
Area of Science:
- Medical Informatics
- Data Mining
- Natural Language Processing
Background:
- MEDLINE database documents require multi-label classification due to multiple category assignments.
- Scalable methods are essential for handling hundreds of thousands of medical documents.
Purpose of the Study:
- To develop a novel, scalable system for automated multi-label classification of MEDLINE documents to MeSH keywords.
- To adapt the ACRI data mining algorithm for effective multi-label classification tasks.
Main Methods:
- Modification of the ACRI algorithm to support multi-label classification.
- Testing five distinct classification configurations and various quality measurement methods.
- Experimental comparison of word reoccurrence-based methods against non-recurrent associative classification.
Main Results:
- The proposed system achieved a macro F1 score of 46%, demonstrating high quality on the challenging MEDLINE dataset.
- The classifier's accuracy reached 90%, calculated as the ratio of true positives and true negatives to total examples.
- Word reoccurrence-based methods showed superiority over non-recurrent associative classification.
Conclusions:
- The developed system offers a high-quality solution for the multi-label classification of MEDLINE documents.
- Different configurations optimize for classifying numerous documents (micro F1) or categories with fewer documents (macro F1).
- A balanced approach optimizing the average of macro and micro F1 provides a trade-off between classification objectives.
Related Concept Videos
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe and...
Classification of Neurotransmitters
Classification of Leukocytes
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
Classification of Systems-II
Drug Classes and Categories
