Related Experiment Video
Updated: Sep 26, 2025

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Why was this cited? Explainable machine learning applied to COVID-19 research literature
Lucie Beranová1, Marcin P Joachimiak2, Tomáš Kliegr3
1Department of Econometrics, Faculty of Informatics and Statistics, VSE Praha, W Churchill sq. 4, Prague, Czech Republic.
Machine learning models predict research citation counts using article content and metadata from the CORD-19 corpus. Findings reveal patterns related to animal hosts and therapeutic targets like dipeptidyl peptidase 4 (DPP4).
Area of Science:
- Bibliometrics and Scientometrics
- Computational Biology
- Machine Learning in Scientific Research
Background:
- Traditional bibliometric factors are insufficient for predicting research citation counts.
- The COVID-19 crisis highlighted the need for rapid analysis of vast biomedical literature.
- Understanding citation patterns can guide research funding and identify impactful studies.
Purpose of the Study:
- To develop machine learning models predicting citation counts using article content and metadata.
- To identify novel patterns and factors influencing research impact beyond traditional bibliometrics.
- To compare the performance and interpretability of various machine learning and rule-learning algorithms.
Main Methods:
- Utilized the CORD-19 corpus comprising biology and medicine research articles.
- Employed advanced machine learning for text understanding: BERT, ConceptNet, Pubtator, ScispaCy.
- Applied explanation algorithms (Random Forest, LIME, Shapley values) and compared black-box models with intrinsically explainable rule-learning models (CORELS, CBA).
Main Results:
- Identified significant patterns, including associations with dipeptidyl peptidase 4 (DPP4), a MERS-CoV receptor.
- Found that articles referencing bats and camels as hosts for coronaviruses received more citations.
- Observed potential citation bias favoring authors with Western-sounding names; TF-IDF and binary word incidence showed equal performance, with the latter offering better interpretability.
Conclusions:
- Machine learning, particularly neural networks, can effectively predict citation counts using article content and metadata.
- Rule-based models, enhanced by semantic entity detection, provide valuable insights into citation drivers.
- Further research should explore citation patterns in phylogenetic contexts and investigate DPP4's role in SARS-CoV-2.
More Related Videos
Related Concept Videos
Steps in Outbreak Investigation
Causality in Epidemiology
Statistical Methods for Analyzing Epidemiological Data
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Single Nucleotide Polymorphisms-SNPs

