Related Experiment Video
Updated: Dec 24, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Integrating image caption information into biomedical document classification in support of biocuration
Xiangying Jiang1, Pengyuan Li1, James Kadin2
1The Computational Biomedicine and Machine Learning Lab, Department of Computer & Information Sciences, University of Delaware, 18 Amstel Ave, Newark, DE 19716, USA.
Automated document classification aids biomedical research by efficiently identifying relevant scientific papers. Incorporating figure captions significantly improves accuracy in this task.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Scientific Literature Analysis
Background:
- Scientific literature growth necessitates automated methods for information retrieval in biomedical research.
- Biocuration tasks, like those in the Gene Expression Database (GXD), are labor-intensive due to the vast number of publications.
- Effective automation of document classification is crucial for efficient curation of biological databases.
Purpose of the Study:
- To develop and evaluate a document classification scheme for identifying topic-relevant papers within large biomedical literature collections.
- To enhance existing meta-classification frameworks by incorporating features from figure captions alongside titles and abstracts.
- To support the biocuration classification task by improving the efficiency and accuracy of identifying relevant scientific articles.
Main Methods:
- A meta-classification framework was adapted to include textual features from titles, abstracts, and figure captions.
- The classifier was trained and tested on a large, imbalanced dataset of approximately 60,000 documents curated by the Gene Expression Database (GXD).
- The dataset spanned documents from 2012-2016, with each document represented by its title, abstract, and figure caption text.
Main Results:
- The developed classifier achieved a precision of 0.698, recall of 0.784, f-measure of 0.738, and Matthews correlation coefficient of 0.711.
- The results demonstrate the framework's effectiveness in handling imbalanced datasets common in biocuration.
- Performance significantly improved when incorporating figure caption information compared to using only titles and abstracts.
Conclusions:
- Figure captions contain substantial information that significantly enhances biomedical document classification and curation.
- The proposed framework offers an effective solution for automated identification of relevant scientific literature, aiding biocurators.
- This approach addresses the challenges posed by the increasing volume of biomedical publications and supports efficient knowledge discovery.
More Related Videos
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Classification of Leukocytes
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
Classification of Epithelial Tissues: Overview
Based on the number of cell layers,...
Imaging Biological Samples with Optical Microscopy
In optical microscopy, the specimen to be viewed is placed on a glass slide and clipped on the stage...
Classification of Systems-II
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:

