Related Experiment Video
Updated: Jun 7, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A strategy for cost-effective large language model use at health system-scale
Eyal Klang1,2, Donald Apakama3,4, Ethan E Abbott5,3,4
1Division of Data-Driven and Digital Medicine, Department of Medicine, Icahn School of Medicine at Mount Sinai, New York, NY, USA. eyal.klang@mountsinai.org.
Large language models (LLMs) can streamline clinical workflows, but scaling them presents challenges. Concatenating up to 50 tasks simultaneously with high-capacity LLMs like Llama-3-70b and GPT-4-turbo-128k offers significant cost savings.
Area of Science:
- Artificial Intelligence in Medicine
- Health Informatics
- Computational Health
Background:
- Large language models (LLMs) show promise for optimizing clinical workflows.
- However, the economic and computational demands of implementing LLMs at a health system scale remain largely unexamined.
- Understanding these limitations is crucial for effective enterprise-level adoption.
Purpose of the Study:
- To evaluate the impact of concatenating multiple clinical notes and tasks on LLM performance under varying computational loads.
- To assess the accuracy and output formatting capabilities of different LLMs.
- To identify the limits of LLM utilization and explore cost-efficiency strategies.
Main Methods:
- Assessed ten LLMs of varying capacities and sizes using real-world patient data.
- Conducted over 300,000 experiments with diverse task sizes and configurations.
- Measured question-answering accuracy and output formatting reliability.
Main Results:
- Model performance degraded as the number of questions and notes increased.
- High-capacity models (Llama-3-70b, GPT-4-turbo-128k) demonstrated resilience, maintaining high accuracy.
- GPT-4-turbo-128k performance declined after 50 tasks with large prompts.
- Concatenation of up to 50 simultaneous tasks proved effective for these models.
- Economic analysis revealed up to a 17-fold cost reduction at 50 tasks.
Conclusions:
- LLMs can effectively concatenate up to 50 simultaneous tasks, particularly high-capacity models, after addressing minor failures.
- This approach significantly reduces costs for enterprise-scale healthcare applications.
- The study delineates LLM utilization limits and highlights pathways for cost-effective deployment in healthcare.
More Related Videos
Related Concept Videos
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Health Literacy
Models of Health Promotion and Illness Prevention I
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...
Methods Of Healthcare Delivery System
Managed Care System:
The managed care system is designed to control the cost while maintaining the quality of care. The patient's care from admission to discharge is planned by the primary care provider or the case manager, also known as the gatekeeper. In a managed care system, the number of care providers is...

