Related Experiment Video
Updated: Sep 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Employing General-Purpose and Biomedical Large Language Models with Advanced Prompt Engineering for
Xinyao Zhang1, Nicole Sonne Heckmann1, Manuela Del Castillo Suero1
1Department of Drug Design and Pharmacology, University of Copenhagen, Jagtvej 160, 2100, Copenhagen, Denmark.
Background:
The potential of large language models (LLMs) to automate and support pharmacoepidemiologic study design is an emerging area of interest, yet their reliability remains insufficiently characterized. General-purpose LLMs often display inaccuracies, while the comparative performance of specialized biomedical LLMs in this domain remains unknown.
Methods:
This study evaluated general-purpose LLMs (GPT-4o and DeepSeek-R1) versus biomedically fine-tuned LLMs (QuantFactory/Bio-Medical-Llama-3-8B-GGUF and Irathernotsay/qwen2-1.5B-medical_qa-Finetune) using 46 protocols (2018-2024) from the HMA-EMA Catalogue and Sentinel System. Performance was assessed across relevance, logic of justification, and ontology-code agreement across multiple coding systems using Least-to-Most (LTM) and Active Prompting strategies.
Results:
GPT-4o and DeepSeek-R1 paired with LTM prompting achieved the highest relevance and logic of justification scores, with GPT-4o-LTM reaching a median relevance score of 4 in 8 of 9 questions for HMA-EMA protocols. Biomedical LLMs showed lower relevance overall and frequently generated insufficient justification. All LLMs demonstrated limited proficiency in ontology-code mapping.
Conclusion:
Off-the-shelf general-purpose LLMs currently offered more reliable support for pharmacoepidemiologic study design than the smaller biomedical LLMs evaluated. In the model-controlled comparison (GPT-4o with Least-to-Most versus Active prompting), prompt strategy did not significantly affect performance (paired p = 0.93), and prompt comparisons are therefore reported descriptively. Because the two model groups also differed in scale and instruction-tuning, this contrast should be interpreted as a comparison of readily deployable options rather than of biomedical specialization alone.
Related Concept Videos
Analysis of Population Pharmacokinetic Data
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Pharmacodynamic Models: Overview
Pharmacogenomics: Identification of New Drug Targets
Impact of Pharmacokinetic–Pharmacodynamic Models: Regulatory Decisions
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.