Related Experiment Video
Updated: May 12, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Empowering open medium-sized generative language models for effective structured search in biomedical systematic
Leandra Budau1, Richard Finney1, Faezeh Ensan1
1Department of Biomedical, Electrical, and Computer Engineering, Toronto Metropolitan University, Toronto, Canada.
International Journal of Medical Informatics
|May 10, 2026
Summary
Fine-tuned open-source models like BioGPT excel at generating Boolean queries for biomedical literature searches, outperforming large commercial models. This offers a cost-effective and scalable solution for systematic reviews.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Information Retrieval
Background:
- Systematic Literature Reviews (SLRs) are crucial for public health and clinical decisions.
- Manual Boolean query generation for SLRs is inefficient and error-prone.
- Existing LLM approaches often use costly commercial models, neglecting open-source alternatives.
Purpose of the Study:
- To develop and evaluate a novel framework for automated Boolean query generation using open-source LLMs.
- To compare the performance of fine-tuned BioGPT and BioT5 against baselines and commercial LLMs.
- To provide a cost-effective and adaptable solution for biomedical literature searching.
Main Methods:
- A three-stage framework using medium-sized, open-source generative models (BioGPT, BioT5).
- Fine-tuning models on PubMed datasets with titles and metadata.
- Evaluating performance on CLEF TAR and FASS-BSLR datasets, comparing against baselines and LLMs.
Main Results:
- Fine-tuned BioGPT surpassed traditional TAR models and commercial LLMs on key retrieval metrics (Precision, F1, MAP, NDCG).
- BioGPT achieved superior recall on the FASS dataset, exceeding GPT-3.5 Turbo, GPT-4, Gemini-2, and Llama-3.
- BioT5 also outperformed most baselines, demonstrating the effectiveness of fine-tuned open-source models.
Conclusions:
- Fine-tuned, open-source, medium-sized generative models are effective for biomedical Boolean query generation.
- These models provide a cost-efficient, privacy-preserving, and scalable alternative to commercial LLMs for SLRs.
- The proposed framework advances automated literature retrieval in biomedical research.
