Related Experiment Video
Updated: Dec 11, 2025

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Target specific mining of COVID-19 scholarly articles using one-class approach
Sanjay Kumar Sonbhadra1, Sonali Agarwal1, P Nagabhushan1
1IIIT Allahabad, Prayagraj, U.P. India 211015.
Insights
Machine learning, specifically k-means clustering and one-class support vector machines (OCSVMs), effectively categorizes COVID-19 research. This approach aids researchers in navigating the vast literature on coronavirus disease 2019 (COVID-19) prevention and treatment.
Area of Science:
- Computational Biology
- Data Science
- Infectious Disease Research
Background:
- The COVID-19 pandemic, caused by SARS-CoV-2, has led to a surge in research publications.
- Manually extracting relevant information from the extensive body of COVID-19 literature is impractical.
- Efficiently identifying research trends and activities is crucial for advancing prevention and treatment strategies.
Purpose of the Study:
- To develop and validate a machine learning approach for analyzing and categorizing COVID-19 research articles.
- To assist the research community in navigating the vast scientific literature on coronavirus disease 2019 (COVID-19).
- To identify trends and activities within COVID-19 research for future exploration of prevention and treatment techniques.
Main Methods:
- Utilized the COVID-19 Open Research Dataset (CORD-19) for experimental analysis.
- Employed clustering techniques, specifically k-means, to group similar research articles.
- Applied parallel one-class support vector machines (OCSVMs) for task assignment and classification of article clusters.
Main Results:
- The combination of k-means clustering followed by parallel OCSVMs demonstrated superior performance in categorizing research articles.
- The proposed method proved effective in both original and reduced feature spaces, validating its robustness.
- The approach successfully mined target-class guided information, revealing patterns in COVID-19 research.
Conclusions:
- The machine learning methodology, particularly k-means clustering with parallel OCSVMs, is an effective tool for analyzing large-scale research datasets like CORD-19.
- This approach facilitates efficient exploration of scientific literature, aiding researchers in identifying key trends and knowledge gaps in COVID-19 research.
- The findings support the use of advanced data analytics for accelerating scientific discovery in response to global health crises.
Abstract:
The novel coronavirus disease 2019 (COVID-19) began as an outbreak from epicentre Wuhan, People's Republic of China in late December 2019, and till June 27, 2020 it caused 9,904,906 infections and 496,866 deaths worldwide. The world health organization (WHO) already declared this disease a pandemic. Researchers from various domains are putting their efforts to curb the spread of coronavirus via means of medical treatment and data analytics. In recent years, several research articles have been published in the field of coronavirus caused diseases like severe acute respiratory syndrome (SARS), middle east respiratory syndrome (MERS) and COVID-19. In the presence of numerous research articles, extracting best-suited articles is time-consuming and manually impractical. The objective of this paper is to extract the activity and trends of coronavirus related research articles using machine learning approaches to help the research community for future exploration concerning COVID-19 prevention and treatment techniques. The COVID-19 open research dataset (CORD-19) is used for experiments, whereas several target-tasks along with explanations are defined for classification, based on domain knowledge. Clustering techniques are used to create the different clusters of available articles, and later the task assignment is performed using parallel one-class support vector machines (OCSVMs). These defined tasks describes the behavior of clusters to accomplish target-class guided mining. Experiments with original and reduced features validate the performance of the approach. It is evident that the k-means clustering algorithm, followed by parallel OCSVMs, outperforms other methods for both original and reduced feature space.
More Related Videos
03:08Using Human Differentially Expressed Gene Lists to Perform Downstream Pathway Enrichment Analysis and Target Prioritization
Published on: October 3, 2025
09:33Author Spotlight: Finding New Therapeutic Targets for Malignant Peripheral Nerve Sheath Tumor Through Genome-Scale shRNA Screens
Published on: August 25, 2023
Related Concept Videos
Targeted Cancer Therapies
There are several types of targeted therapies against...
Single Nucleotide Polymorphisms-SNPs
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Dose-Response Relationship: Selectivity and Specificity
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Chi-square Analysis
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...