Related Experiment Video
Updated: May 21, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Screening nonrandomized studies for medical systematic reviews: a comparative study of classifiers
Tanja Bekhuis1, Dina Demner-Fushman
1Department of Biomedical Informatics, School of Medicine, University of Pittsburgh, Pittsburgh, PA 15232, USA. tcb24@pitt.edu
Machine learning classifiers can identify nonrandomized studies for systematic reviews. Optimization significantly improves classifier performance, substantially reducing the number of citations requiring reviewer screening.
Area of Science:
- Biomedical Informatics
- Systematic Review Methodology
- Machine Learning Applications
Background:
- Systematic reviews often include nonrandomized studies.
- Screening citations for full-text review is time-consuming.
- Automated methods could improve efficiency.
Purpose of the Study:
- Evaluate machine learning (ML) classifiers for identifying nonrandomized studies.
- Assess the impact of optimization on classifier performance.
- Determine the potential reduction in screening workload.
Main Methods:
- Utilized an open-source data-mining suite for citation classification.
- Trained and tested ML classifiers (k-nearest neighbor, naïve Bayes, evolutionary support vector machine) on biomedical citations.
- Compared performance across different feature sets (bag of words, n-grams) and citation portions (titles, abstracts, full citations).
- Investigated optimization techniques including manual thresholding and cross-validation.
Main Results:
- No classifier achieved sufficient recall without optimization in Phase I.
- Optimization significantly improved performance for evolutionary support vector machine and complement naïve Bayes classifiers in Phase II.
- Evolutionary support vector machine and complement naïve Bayes classifiers reduced the initial retrieval set by 46% and 35%, respectively.
- Generalization performance varied among classifiers.
Conclusions:
- Machine learning classifiers are effective tools for identifying nonrandomized studies for systematic reviews.
- Optimization is crucial for enhancing classifier performance.
- Classifier generalizability differs, but substantial reductions in screening workload are achievable.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which result in visible changes...
Randomized Experiments
Simple randomization
Simple...
Hazard Ratio
For example, in a clinical trial evaluating a...
Bioequivalence Experimental Study Designs: Completely Randomized and Randomized Block Designs
