Related Experiment Video
Updated: Mar 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation of a Retrieval-Augmented Generation Chatbot for Antimicrobial Resistance Research: Comparative Analysis of
Oscar Escudero-Arnanz1, Manuel Eduardo Valero-Méndez1, Noelia Sánchez-Ramos1
1Department of Signal Theory and Communications, King Juan Carlos University, Camino del molino 5, Fuenlabrada, Madrid, 28943, Spain, 34689582328.
Background:
Antimicrobial resistance (AMR) poses a critical global health threat, undermining the efficacy of antibiotics and complicating clinical decision-making. Although scientific literature on AMR is extensive, retrieving and synthesizing relevant evidence remains time-consuming for clinicians and researchers. Recent advances in large language models (LLMs) offer opportunities to enhance access to domain-specific knowledge. However, the diversity of available models, ranging from open-source to commercial, necessitates a systematic comparison of their performance, cost, and scalability in real-world biomedical applications.
Objective:
This study aims to describe the development of a retrieval-augmented generation (RAG) chatbot for AMR literature analysis and compare multiple commercial and open-source LLMs in terms of accuracy, faithfulness, response time, and cost-efficiency.
Methods:
A corpus of 164 peer-reviewed AMR-related articles was compiled from Google Scholar and embedded into a ChromaDB vector database using OpenAI's text-embedding-ada-002. The RAG chatbot was implemented to operate with 5 LLM backbones: GPT-4, GPT-4o, GPT-4o-mini, Claude 3.7 Sonnet, and LLaMA 4 Maverick. For each model, a temperature ablation study was performed to determine optimal performance. Evaluation metrics included correctness (pass rate and score), faithfulness, relevancy, computational cost, and latency, using a synthetic ground truth dataset generated with GPT-4.
Results:
All models generated scientifically grounded responses when integrated into the RAG framework. GPT-4 achieved the highest correctness score (94.7%) but incurred the highest cost, while GPT-4o delivered nearly identical accuracy at a 9-fold lower cost and the fastest response time (3.88 s). LLaMA 4 Maverick and GPT-4o-mini offered lower accuracy but substantially reduced operational costs. Claude 3.7 Sonnet showed competitive accuracy, but the least favorable cost-performance ratio. Qualitative analysis revealed differences in response style, detail, and structure among models.
Conclusions:
A RAG-based chatbot can effectively support AMR research by delivering accurate, context-grounded, and scalable access to scientific literature. The comparative evaluation highlights trade-offs between performance, cost, and speed, guiding the selection of LLM architectures for clinical and research settings. Future work will focus on integrating language-specific embeddings and specialized domain agents to further enhance accuracy, adaptability, and clinical use.
Insights
A new retrieval-augmented generation (RAG) chatbot effectively analyzes antimicrobial resistance (AMR) literature. GPT-4o offers a cost-effective and fast solution, balancing accuracy and performance for researchers.
Area of Science:
- Biomedical Informatics
- Artificial Intelligence in Medicine
- Computational Biology
Background:
- Antimicrobial resistance (AMR) is a significant global health challenge, impacting antibiotic efficacy and clinical decisions.
- Synthesizing extensive AMR literature is time-consuming for researchers and clinicians.
- Large language models (LLMs) present opportunities to improve access to specialized biomedical knowledge.
Purpose of the Study:
- Develop a retrieval-augmented generation (RAG) chatbot for analyzing antimicrobial resistance (AMR) literature.
- Compare the performance, cost, and scalability of various commercial and open-source LLMs for AMR research.
Main Methods:
- Compiled a corpus of 164 AMR articles and embedded them into a ChromaDB vector database.
- Implemented a RAG chatbot utilizing five LLM backbones: GPT-4, GPT-4o, GPT-4o-mini, Claude 3.7 Sonnet, and LLaMA 4 Maverick.
- Evaluated models based on correctness, faithfulness, relevancy, computational cost, and latency using a synthetic dataset.
Main Results:
- All tested LLMs produced scientifically accurate responses within the RAG framework.
- GPT-4o achieved high accuracy at a significantly lower cost and faster response time compared to GPT-4.
- LLaMA 4 Maverick and GPT-4o-mini provided cost savings with reduced accuracy, while Claude 3.7 Sonnet had a less favorable cost-performance ratio.
Conclusions:
- RAG-based chatbots can enhance AMR research by providing accurate and scalable access to scientific literature.
- The study highlights the performance-cost-speed trade-offs among different LLMs, aiding selection for clinical and research applications.
- Future work will explore language-specific embeddings and domain agents to further improve chatbot accuracy and clinical utility.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:14Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Related Concept Videos
Antibiotic Selection
Development of Antibiotic Resistance
Clinical Significance of Antibiotic Resistance
Antimicrobial Proteins
Interferons
Interferons (IFNs) are proteins produced by lymphocytes, macrophages, and fibroblasts infected with viruses. While IFNs cannot prevent viruses from entering and...
Mechanism of Antibiotic Resistance in MRSA
Antimicrobial Effectiveness