Related Experiment Video
Updated: Aug 30, 2026

Bringing the Clinic Home: An At-Home Multi-Modal Data Collection Ecosystem to Support Adaptive Deep Brain Stimulation
Published on: July 14, 2023
Clinician-led, AI-assisted clinical software development: multi-domain verification of a deployed ePROM application
Alberto Zambudio Munuera1, Miguel Ángel Arrabal Polo1, Clemente García Hidalgo2
1Department of Urology, Hospital Universitario Clínico San Cecilio, Granada, Spain.
Background:
Generative AI lowers the technical barrier to clinician-led software development, but functional success and usability do not establish clinical or technical assurance.
Objective:
To characterise defects that survived deployment in a clinician-built, AI-assisted ePROM application and the assurance activities that detected them.
Methods:
Single-case retrospective development-and-assurance report of STUIapp, a browser-accessible application integrating six validated lower urinary tract symptom instruments. Assurance comprised structured post-hoc code audit, independent clinical review of a frozen 78-case scoring matrix, WCAG 2.1 parameter measurement, a 23-canary persistent-storage study, formative usability evaluation with 14 clinicians, 26 patients and 12 older adults, preliminary MDR positioning, and targeted post-hoc requirements traceability across 926 attributable author messages with independent classification. Descriptive statistics only.
Results:
Mean SUS was 92.3 (SD 8.8) in clinicians and 92.0 (SD 10.8) in patients. Although all 78 original automated cases passed, independent review found 17 expected results requiring clinical or methodological correction. Technical findings included instrument mislabelling, a manifest that failed installability validation, and a third-party analytics tag contradicting the local-only privacy claim. Accessibility measurement found 10-px text and 2.56:1 contrast. Persistent-storage permission was denied in all browser-tab canaries (10/10) and granted in all installed canaries (13/13); no eviction was observed through 38.3 days. Home-screen addition was unaided in 7/12 older adults and 0/4 with limited smartphone use. Within selected defect classes, requirement traceability showed heterogeneous failure patterns.
Conclusions:
Usability and functional success did not establish clinical correctness, privacy, standards-based accessibility or regulatory preparedness. These findings support independent, domain-specific assurance of both implementation and the clinical expectations and non-functional requirements encoded in clinician-led AI-assisted clinical software.
