Related Experiment Video
Updated: May 16, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative Evaluation of Artificial Intelligence Large Language Models for Breast Reconstruction Patient Education
Berk B Ozmen1, Victor F A Almeida2, Ibrahim Berber3
1Department of Plastic Surgery, Cleveland Clinic, Cleveland, OH, USA.
Plastic Surgery (Oakville, Ont.)
|May 15, 2026
Summary
Artificial intelligence (AI) models show varied performance in breast reconstruction patient education. General large language models (LLMs) offer broad utility, while specialized retrieval-augmented generation (RAG) systems excel in specific clinical recovery topics.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Surgery
- Patient Education
Background:
- Informed decision-making in breast reconstruction surgery relies heavily on patient education.
- Large language models (LLMs) present potential for medical information dissemination, but their accuracy in specialized surgical fields requires evaluation.
Purpose of the Study:
- To assess the performance of general-purpose LLMs and a specialized retrieval-augmented generation (RAG) system in providing patient education for breast reconstruction.
- To compare the accuracy, relevance, clarity, and completeness of AI-generated responses.
Main Methods:
- Ten standardized questions on breast reconstruction were posed to five AI systems: ChatGPT o3-high, ChatGPT 4.5, Grok 3, Claude Haiku 3.5, and a specialized MicroRAG system.
- Responses were evaluated by four plastic surgeons using a Global Quality Score (1-5 scale).
Main Results:
- Performance varied, with ChatGPT o3-high scoring highest overall (3.73).
- The specialized MicroRAG system achieved perfect scores for clinical recovery topics, offering evidence-based, cited responses.
- ChatGPT o3-high significantly outperformed ChatGPT 4.5 (P=.005).
Conclusions:
- AI systems exhibit complementary strengths for breast reconstruction patient education.
- General LLMs provide consistent information across various patient needs.
- Specialized RAG systems offer superior, evidence-based answers in specific clinical areas, suggesting a need for tailored AI tool selection.