Related Experiment Video
Updated: Sep 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Statistical Analysis of the Performance of Open-Access Language Models: A Tool in the Clinical Laboratory
Micaela Avellaneda1, Pamela De Francesco2, Nilda Fink3
1B.Sc. in Biochemistry Federacion Bioquimica de la Provincia de Buenos Aires.
Introduction:
Artificial intelligence (AI) has been progressively incorporated into the clinical analysis laboratory, optimizing data management, detection of complex patterns, and image-based diagnostic support. Among emerging tools, Large Language Models (LLMs) represent a significant advance by enabling natural language processing for clinical, analytical, and academic tasks. However, their performance is heterogeneous and depends on model architecture, training, and contextual fine-tuning. Objective: To compare the performance of eight openly accessible large language models (LLMs) applied to the biomedical domain by assessing their output quality and inter-rater consistency.
Materials And Methods:
A cross-sectional comparative study was carried out. Eight LLMs (ChatGPT-4, DeepSeek V3.1, Gemini, Copilot GPT-5, Claude 3.5 Haiku, Perplexity AI, DeepAI, and Poe) were queried using six standardized questions regarding the analytical evolution of blood glucose measurement. Responses were evaluated by two independent automated systems (Text Cortex AI and Grok Code Fast 1) using a Likert scale from 1 to 10 and considering eight criteria of textual and scientific quality. Friedman and Wilcoxon-Holm tests were used for global and pairwise comparisons, and Pearson and Spearman coefficients were calculated to assess inter-rater agreement.
Results:
Significant differences were identified among models (p < 0.001). The LLMs Gemini, DeepSeek, and ChatGPT-4 obtained the highest scores (medians > 8.5), whereas DeepAI and Poe showed lower performance. Inter-rater correlation based on median scores showed a markedly strong linear association (r = 0.91; p = 0.0019) and a robust ordinal agreement (ρ = 0.83; p = 0.011). The least-squares regression line (E2 = 1.37·E1 - 3.25) indicated a sustained positive relationship between the two systems, with balanced dispersion around the linear model. These findings confirm the stability of the evaluation pattern and support the reproducibility of the method.
Conclusion:
The LLMs with greater complexity and contextual fine-tuning demonstrated superior accuracy, coherence, and reproducibility; however, it is suggested that LLMs may serve as complementary tools depending on the clinical or research context. Their use in healthcare requires methodological validation and expert oversight to ensure safety and clinical applicability within established ethical and regulatory frameworks.
Related Concept Videos
Statistical Software for Data Analysis and Clinical Trials
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
