Related Experiment Video
Updated: Jun 5, 2026

Drug Repurposing Hypothesis Generation Using the "RE:fine Drugs" System
Published on: December 11, 2016
Optimising screening efficiency in evidence synthesis on health Technology: A simulation study using ASReview.
Júlia Meller Dias de Oliveira1, Arthur Thives Mello2, Daniel Henrique Scandolara3
1Bridge Laboratory, Federal University of Santa Catarina, Florianópolis, Brazil; Graduate Program in Dentistry, Federal University of Santa Catarina, Florianópolis, Brazil.
Active learning screening for health technology reviews effectively reduces workload. The Support Vector Machine with Term Frequency-Inverse Document Frequency (SVM + TF-IDF) model and a 7% consecutive-irrelevant stopping rule showed strong performance, but optimal decisions require considering review specifics.
Area of Science:
- Health Technology Assessment
- Systematic Reviews
- Evidence Synthesis
- Machine Learning in Healthcare
Background:
- Active learning (AL) streamlines title-and-abstract screening in health technology evidence syntheses.
- Optimizing AL model configuration and stopping rules is crucial for efficient and reliable screening.
- Previous research has not exhaustively evaluated various AL configurations and stopping rules in this context.
Purpose of the Study:
- To evaluate model-configuration and stopping-rule decisions for active learning-based title-and-abstract screening.
- To identify optimal AL strategies for health technology evidence syntheses.
- To assess the impact of different classifiers and feature extraction methods on screening performance.
Main Methods:
- Retrospective simulations using seven pre-labelled health technology review datasets.
- Comparison of lightweight AL configurations (one-hot encoding, TF-IDF) with multiple classifiers (Naive Bayes, Logistic Regression, Random Forest, SVM).
- Performance evaluation using Normalized Recall Regret (loss), Work Saved over Sampling (WSS), early recall, and K%-consecutive-irrelevant stopping rules.
Main Results:
- Support Vector Machine (SVM) with Term Frequency-Inverse Document Frequency (TF-IDF) using bigrams demonstrated superior performance (average loss: 0.08, WSS@95: 0.70).
- A fixed 7% consecutive-irrelevant stopping rule achieved high recall (mean 98%) across most datasets.
- Dataset characteristics like relevant-record prevalence and textual similarity influenced model performance and stopping-rule reliability.
Conclusions:
- Active learning significantly reduces workload in health technology evidence syntheses.
- SVM + TF-IDF (with bigrams) is a pragmatic initial AL configuration.
- A 7% consecutive-irrelevant stopping rule is a useful heuristic, but final stopping decisions should integrate review-specific factors beyond a fixed threshold.
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Hazard Ratio
For example, in a clinical trial evaluating a...
Study Designs in Epidemiology
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and case-control studies.
Comparing the Survival Analysis of Two or More Groups
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Clinical Trials
There are four phases in a clinical trial. A phase one...