Related Experiment Video
Updated: Oct 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Entity-centric evaluation of large language model responses for medical question-answering tasks
Yi Liu1, Vijaya B Kolachalama1,2,3
1Faculty of Computing & Data Sciences, Boston University, Boston, Massachusetts, United States of America.
Abstract:
Large language models (LLMs) are increasingly evaluated for clinical question-answering (QA) tasks, yet benchmark accuracy alone provides limited assurance of clinical reliability. Existing evaluation metrics often fail to capture whether model reasoning preserves patient-specific context and diagnostic intent, while many require external references or manual annotation, limiting scalability and real-world applicability. To address this gap, this study proposes EntQA, a reference-free, entity-centric metric that evaluates how well LLM-generated responses retain clinically relevant biomedical concepts from patient backgrounds and diagnostic questions. Across five medical QA benchmarks and seven Qwen 2.5 Instruct models (0.5B-72B parameters), EntQA showed consistently positive associations with model accuracy and scaling, outperforming conventional overlap- and embedding-based metrics that frequently exhibited weak or negative correlations. Group-level correlations with accuracy reached Spearman ρ =0.9286, while correlations with model scale reached ρ = 0.252. These findings suggest that EntQA provides a scalable and interpretable framework for assessing clinical fidelity and reasoning quality in healthcare LLMs without requiring gold standard references or external evidence.