Related Experiment Video
Updated: Mar 7, 2026

Drug Repurposing Hypothesis Generation Using the "RE:fine Drugs" System
Published on: December 11, 2016
TrialSieve: A Comprehensive Biomedical Information Extraction Framework for PICO, Meta-Analysis, and Drug Repurposing
David Kartchner1, Haydn Turner1, Christophe Ye1
1Laboratory for Pathology Dynamics, Georgia Institute of Technology, Emory University School of Medicine, Atlanta, GA 30332, USA.
TrialSieve enhances biomedical information extraction for clinical meta-analysis and drug repurposing. Automated NLP models trained on its data can match or exceed human performance in annotation tasks.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Clinical Research
Background:
- Traditional PICO (Patient, Intervention, Comparison, Outcome) methods lack quantitative comparison capabilities for clinical outcomes.
- Biomedical information extraction is crucial for meta-analysis and drug repurposing but faces challenges with data complexity.
- Existing annotation frameworks may not fully capture the nuances required for comprehensive systematic reviews.
Purpose of the Study:
- Introduce TrialSieve, a novel framework for biomedical information extraction.
- Enhance clinical meta-analysis and drug repurposing through improved data annotation and comparison.
- Evaluate the performance of various NLP models and a large language model (LLM) using the TrialSieve dataset.
Main Methods:
- Developed TrialSieve, incorporating hierarchical, treatment group-based graphs extending PICO.
- Annotated 1609 PubMed abstracts with 20 categories, resulting in 170,557 annotations and 52,638 spans.
- Evaluated NLP models (BioLinkBERT, BioBERT, KRISSBERT, PubMedBERT) and GPT-4o on the TrialSieve dataset for entity labeling.
- Conducted an annotator user study (n=39) to assess the efficiency and accuracy of the TrialSieve annotation approach.
Main Results:
- BioLinkBERT achieved the highest accuracy (0.875) and recall (0.679) in biomedical entity labeling.
- PubMedBERT demonstrated the best precision (0.614) and F1-score (0.639).
- NLP models trained on imperfectly annotated data matched or surpassed human performance, indicating feasibility of automation.
- The TrialSieve tree-based approach significantly improved annotator efficiency and accuracy (p < 0.05).
Conclusions:
- TrialSieve offers a robust foundation for automated biomedical information extraction.
- The framework facilitates more comprehensive and quantitative comparisons for clinical meta-analysis and drug repurposing.
- Automated information extraction using NLP models is feasible even with noisy, human-annotated datasets.
More Related Videos
Related Concept Videos
Drug Discovery: Overview
Clinical Trials: Overview
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Statistical Software for Data Analysis and Clinical Trials
Drug Biotransformation: Overview
Pharmacogenomics: Identification of New Drug Targets

