Related Experiment Video
Updated: Jun 30, 2026

Occlusion of the Great and Small Saphenous Vein Using Copolymeric Glue Based on N-Butyl Cyanoacrylate and Methacryloxy Sulfolane
Published on: December 9, 2022
Benchmarking public large language model responses to patient-facing varicose veins questions: informational quality,
Wei Zhong1, Guoxue Zheng1, Qin Li1
1Department of Vascular Surgery, Suining Central Hospital, Suining, China.
Objectives:
To benchmark the informational quality, verifiability indicators, and readability of publicly accessible large language model (LLM) responses to standardized, patient-facing varicose veins (VV) questions.
Methods:
Twenty single-intent VV questions were derived from PubMed-indexed VV and chronic venous disease guidelines and consensus statements (search date: February 10, 2026). The question set was designed as a decision-critical benchmark across the care pathway rather than a prevalence-weighted sample of real-world patient queries. Five publicly accessible LLMs (ChatGPT 5.2, DeepSeek-V3.2, Gemini 3 Pro, Grok 4.1, and Qwen3-Max) were queried through their official web interfaces from February 10 to 12, 2026, under default settings, generating 100 responses. Each prompt was entered in a new privacy-mode session, and refusals or other non-responsive outputs were retained as returned. Two blinded clinicians independently rated DISCERN (16-80), EQIP (0%-100%), GQS (1-5), and the JAMA benchmark (0-4). In this study, JAMA was used as a structured measure of visible attribution and verifiability-related features rather than as a comprehensive measure of transparency for conversational AI. Readability was assessed using six standard indices. Interrater reliability was evaluated with ICC(A,1) and weighted Cohen's κ. Between-model differences were tested using Friedman tests with Kendall's W and Holm adjustment.
Results:
Interrater reliability was high [DISCERN ICC(A,1) = 0.913; EQIP ICC(A,1) = 0.859; GQS κ = 0.883; JAMA κ = 0.864]. Informational-quality scores were broadly similar across models (DISCERN means, 46.50-50.75; EQIP means, 71.50-74.25; GQS medians, 4.0). JAMA scores were uniformly low (means, 0.00-0.25; medians, 0), indicating sparse visible attribution and limited verifiability cues in default outputs. Between-model differences in the primary informational-quality outcomes were small and were not significant after Holm adjustment. Readability differences were more pronounced, and all models exceeded commonly recommended sixth-grade readability thresholds.
Conclusions:
Under default public-user settings, publicly accessible LLMs generated fluent VV responses with limited visible verifiability indicators and suboptimal readability. Differences in the primary informational-quality outcomes were modest and should be interpreted cautiously. This benchmark evaluates communication-related performance rather than claim-level clinical accuracy or safety. These findings support efforts to improve auditability, provenance reporting, and uncertainty communication, but these dimensions do not substitute for formal assessment of factual accuracy, guideline concordance, and clinical safety.
Related Concept Videos
Varicose Veins II: Diagnostic Studies and Interprofessional Care
Assessing Blood pressure in the Leg
Preparation:
Assessment of the Cardiovascular System III: Palpation
Jugular Venous Pressure (JVP) Measurement
Position the patient at a thirty- to forty-five-degree angle or in a semi-fowler's position. Look for the highest point of pulsation in the internal jugular vein and measure the vertical distance to the angle of Loius or sternal angle. A normal JVP is 3-4 cm above the...
Varicose Veins I: Introduction
Veins of Lower Limbs
Formed by the union of the medial and lateral plantar veins, the posterior tibial vein, rising through the calf muscle, assimilates the fibular vein. The anterior tibial vein, a superior extension of the foot's dorsalis pedis vein, merges with the posterior tibial vein at the knee,...
Venous Thrombosis II: Clinical Manifestations and Diagnostic Studies