Related Experiment Video
Updated: Sep 25, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Development of a Framework for Evaluating Large Language Model Safety and Reliability: a Proof-of-Concept Evaluation
Fangyan Liu1,2, Zhi Liu1,2, Xiaolu Fei3
1Department of Emergency, National Clinical Research Center for Geriatric Diseases, Xuanwu Hospital of Capital Medical University, Beijing, 100053, China.
Abstract:
Large language models (LLMs) are entering clinical decision support faster than methodology can characterise their safety. Aggregate accuracy treats all errors as interchangeable and cannot support safe deployment under Software as a Medical Device (SaMD) and EU AI Act frameworks. To develop and demonstrate a framework for evaluating large language model safety and reliability using catastrophic-failure frequency and response reproducibility, incorporating a pre-specified error taxonomy, difficulty-stratified analysis, and two-layer response consistency. The framework was applied to 54 diagnostically challenging emergency cases from non-public institutional records. Six LLMs were accessed via application programming interfaces (APIs) (text-only, zero-shot, defaults; August-December 2025) and queried three times each, producing 972 physician-scored responses. Aggregate scores concealed heterogeneity: anchoring-bias failure (score ≤ 3) varied nearly four-fold (6.9%-26.4%) and rare-disease recognition six-fold (6.1%-37.9%); catastrophic (dangerous-recommendation, score ≤ 2) rates were two-tiered (1.5% for the safest two vs. 6.0% for the rest; p < 0.001), though within-tier ranks were inseparable (2-12 per model). Despite a middle-tier mean, GPT-5.1 was least reproducible (within-case SD 2.69) and fell within the higher catastrophic-failure tier; across models, 31 of 34 dangerous combinations were stochastic, not systematic-a failure mode hidden by aggregate metrics. Open-source DeepSeek R1 was statistically equivalent to Gemini 3 Pro within a 1.5-point margin by two one-sided tests (TOST: Δ - 0.09; 90% CI - 0.47 to 0.30; p < 0.001). In this single-centre proof-of-concept, aggregate accuracy was insufficient for safety characterisation: models with indistinguishable mean accuracy carried different catastrophic-failure tiers and reproducibility profiles. Error-type and response-consistency profiling may inform model selection, ensemble design, and conformity assessment.