Related Experiment Video
Updated: Jun 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Clinical Quality-Assurance Panel for Patient-Facing Black-Box Large Language Models in Healthcare: Anchored
1Graduate School of Public Policy, Hosei University, Tokyo, JPN.
Abstract:
Background Patient-facing large language models (LLMs) are being integrated into patient portals, automated health messaging, psychoeducation, and mental health-related support. Because many deployments are vendor-hosted black boxes that may change without local visibility, health systems need practical quality-assurance methods that can establish a baseline behavioral signature and flag candidate changes in how sensitive topics are framed. Methods This feasibility study applied an anchored two-alternative forced-choice (2AFC) protocol to ChatGPT 5.4 (OpenAI, San Francisco, California, United States) through its consumer-facing interface on 25 March 2026. The primary objective was to determine whether a version-controlled anchored 2AFC panel could produce a parsable, repeatable, clinically interpretable baseline signature for a black-box LLM under a fixed prompt family. Secondary objectives were to estimate an anchored scale, place selected clinically salient concepts on that scale, quantify repeatability and confidence behavior, and explore sensitivity to wrapper and prompt-family changes. On each of 190 trials, the model compared two same-sized "solids" and chose which was harder (resistance to indentation and scratching; brittleness excluded), then returned a one-line JavaScript Object Notation (JSON) object containing a forced-choice response and a confidence score (0-100). A Bradley-Terry model generated an anchored 0-100 hardness scale (Marshmallow = 0; Steel = 100) with parametric bootstrap confidence intervals (CIs). We also executed the same panel on a locally hosted model to assess on-premise feasibility. Results The anchored panel produced a coherent within-session baseline signature with complete repeated-pair agreement (all 64 repeated unordered pairs yielded identical winners) and near-perfect Bradley-Terry fit (classification accuracy 1.000; log loss 0.0117). The scale separated lower-hardness concepts (Life 6.04, 95% CI -0.79 to 8.69; Kindness 13.70, 95% CI 4.42 to 16.79; Silence 21.14, 95% CI 12.34 to 24.09) from higher-hardness concepts (Justice 75.93, 95% CI 71.59 to 83.91; Death 87.15, 95% CI 84.59 to 88.31). Model-reported confidence correlated strongly with inferred pair distance (r = 0.842, 95% CI 0.813-0.869). Style/language wrappers showed complete agreement, whereas a prompt-family comparison produced three reversals among 36 hardness pairs. In the local-model feasibility run, internal consistency was lower (accuracy 0.926; log loss 0.158) with anchor-order violations (Polycarbonate sheet > Steel), demonstrating that the panel can flag deviations that warrant clinical or governance review. Conclusions Anchored forced-choice panels can serve as compact, repeatable behavioral control materials for patient-facing LLM governance by establishing a baseline for later regression testing, post-update monitoring, and prompt-template change control. This single-model, single-session feasibility study does not demonstrate longitudinal drift detection, clinical outcomes, or therapeutic appropriateness. The instrument should therefore complement, not replace, scenario-based clinical safety evaluation, expert review, and organizational oversight.