Related Experiment Video
Updated: Aug 30, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Multi-label classification for biomedical literature: an overview of the BioCreative VII LitCovid Track for COVID-19
Qingyu Chen1, Alexis Allot1, Robert Leaman1
1National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, MD, Bethesda 20892, USA.
Insights
Automated topic annotation for COVID-19 literature was developed using a new large dataset. This approach significantly improves upon existing methods for classifying research articles, aiding information discovery.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Computational Biology
Background:
- The COVID-19 pandemic generated a massive volume of biomedical literature, overwhelming manual curation efforts.
- Accurate topic annotation of COVID-19 research is crucial for navigating and utilizing this rapidly expanding knowledge base.
- Existing text-mining methods have not adequately addressed the specific challenge of topic annotation in this domain.
Purpose of the Study:
- To address the bottleneck in manual curation of COVID-19 literature by developing automated topic annotation methods.
- To establish a benchmark for automated topic annotation through the BioCreative LitCovid track.
- To create and release a large-scale, multi-label dataset for training and evaluating topic annotation models.
Main Methods:
- Organized the BioCreative LitCovid track, a community effort for automated topic annotation.
- Created the BioCreative LitCovid dataset, comprising over 30,000 manually reviewed COVID-19 articles.
- Evaluated 80 submissions from 19 international teams, primarily using transformer-based hybrid systems.
Main Results:
- The highest-performing systems achieved macro-F1 scores of 0.8875, micro-F1 scores of 0.9181, and instance-based F1 scores of 0.9394.
- These results substantially surpassed the performance of state-of-the-art multi-label classification methods, demonstrating significant improvement.
- The BioCreative LitCovid track successfully fostered advancements in automated topic annotation for biomedical literature.
Conclusions:
- Automated topic annotation for COVID-19 literature is feasible and highly effective with advanced methods.
- The developed dataset and track provide a valuable resource for future research in biomedical text mining.
- This work significantly closes the gap between dataset curation and method development in managing pandemic-related scientific information.
Abstract:
The coronavirus disease 2019 (COVID-19) pandemic has been severely impacting global society since December 2019. The related findings such as vaccine and drug development have been reported in biomedical literature-at a rate of about 10 000 articles on COVID-19 per month. Such rapid growth significantly challenges manual curation and interpretation. For instance, LitCovid is a literature database of COVID-19-related articles in PubMed, which has accumulated more than 200 000 articles with millions of accesses each month by users worldwide. One primary curation task is to assign up to eight topics (e.g. Diagnosis and Treatment) to the articles in LitCovid. The annotated topics have been widely used for navigating the COVID literature, rapidly locating articles of interest and other downstream studies. However, annotating the topics has been the bottleneck of manual curation. Despite the continuing advances in biomedical text-mining methods, few have been dedicated to topic annotations in COVID-19 literature. To close the gap, we organized the BioCreative LitCovid track to call for a community effort to tackle automated topic annotation for COVID-19 literature. The BioCreative LitCovid dataset-consisting of over 30 000 articles with manually reviewed topics-was created for training and testing. It is one of the largest multi-label classification datasets in biomedical scientific literature. Nineteen teams worldwide participated and made 80 submissions in total. Most teams used hybrid systems based on transformers. The highest performing submissions achieved 0.8875, 0.9181 and 0.9394 for macro-F1-score, micro-F1-score and instance-based F1-score, respectively. Notably, these scores are substantially higher (e.g. 12%, higher for macro F1-score) than the corresponding scores of the state-of-art multi-label classification method. The level of participation and results demonstrate a successful track and help close the gap between dataset curation and method development. The dataset is publicly available via https://ftp.ncbi.nlm.nih.gov/pub/lu/LitCovid/biocreative/ for benchmarking and further development. Database URL https://ftp.ncbi.nlm.nih.gov/pub/lu/LitCovid/biocreative/.
More Related Videos
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Genomics
Classification of Leukocytes
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Clinical Trials: Overview
Cardiovascular Drugs: Classification based on Therapeutic Indications