Related Experiment Video
Updated: Mar 11, 2026

Eye-tracking Technology and Data-mining Techniques used for a Behavioral Analysis of Adults engaged in Learning Processes
Published on: June 10, 2021
Interactive active learning for literature screening: finetuning GPT with DeepSeek reasoning for cross-domain
Yiming Li1,2, Joseph M Plasek1,2, Xinsong Du1,2
1Department of Medicine, Harvard Medical School, Boston, MA 02115, United States.
This study introduces an active learning framework that uses disagreement between large language models (LLMs) to improve biomedical literature screening. Fine-tuning GPT models with this method significantly boosted performance, especially for lighter models like GPT-4o-mini.
Area of Science:
- Biomedical Informatics
- Artificial Intelligence in Medicine
Background:
- Automated literature screening in biomedicine faces challenges from domain shifts and limited labeled data.
- Large language models (LLMs) struggle with complex, domain-specific reasoning in zero-shot settings.
Purpose of the Study:
- To investigate an interactive, weakly supervised learning framework combining GPT's fine-tuning with DeepSeek's reasoning for improved biomedical literature screening.
- To enhance model accuracy and generalizability across diverse biomedical domains.
Main Methods:
- Developed an active learning framework using model disagreement between GPT-4o and DeepSeek to identify misclassified articles.
- Fine-tuned three GPT variants (GPT-4o, GPT-4o-mini, GPT-4.1-nano) using disagreement-based samples and DeepSeek's rationale traces for weak supervision.
- Evaluated performance on independent benchmark sets in cancer immunotherapy and LLMs in medicine, prioritizing recall.
Main Results:
- Fine-tuning GPT models with disagreement-based examples significantly improved performance.
- GPT-4o-mini achieved the highest F1 score (0.93) and recall (0.95) after fine-tuning.
- Fine-tuned models consistently outperformed zero-shot counterparts across biomedical topics without increasing reviewer workload.
Conclusions:
- Disagreement-driven active learning effectively enhances GPT-based biomedical literature screening.
- Lightweight models like GPT-4o-mini show significant benefits from targeted, reasoning-enriched training.
- The framework offers a scalable solution for efficient and reliable information retrieval in systematic reviews.
Related Concept Videos
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Purposive Learning
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Critical Thinking II
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Deductive Reasoning
For example, a researcher can deduce specific predictions...
