Related Experiment Video
Updated: May 22, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking Large Language Models from Open and Closed Source Models to Apply Data Annotation for Free-Text Criteria
Ali Nemati1, Mohammad Assadi Shalmani1, Qiang Lu2
1Health Informatics Department, Zilber College of Public Health, University of Wisconsin, Milwaukee, WI 53211, USA.
None:
Large language models (LLMs) hold the potential to significantly enhance data annotation for free-text healthcare records. However, ensuring their accuracy and reliability is critical, especially in clinical research applications requiring the extraction of patient characteristics. This study introduces a novel evaluation framework based on Multi-Criteria Decision Analysis (MCDA) and the Order of Preference by Similarity to Ideal Solution (TOPSIS) technique, designed to benchmark LLMs on their annotation quality. The framework defines ten evaluation metrics across key criteria such as age, gender, BMI, disease presence, and blood markers (e.g., white blood count and platelets). Using this methodology, we assessed leading open source and commercial LLMs, achieving accuracy scores of 0.59, 1, 0.84, 0.56, and 0.92, respectively, for the specified criteria. Our work not only provides a rigorous framework for evaluating LLM capabilities in healthcare data annotation but also highlights their current performance limitations and strengths. By offering a comprehensive benchmarking approach, we aim to support responsible adoption and decision-making in healthcare applications.
