Related Experiment Video
Updated: Aug 16, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Effects of a Safety User Interface Bundle on Verification Intentions in Generative AI Chat Use Among Older Chinese
Jun'an Yu1, Jun Chen2, Anjie Ren3
1Faculty of Science, University of Auckland, Auckland, New Zealand.
Background:
Generative AI chat systems are increasingly used for everyday information seeking, but plausible errors and omissions can mislead users when outputs are accepted without scrutiny. Interface-level safety cues may help users calibrate trust and engage in verification; yet, evidence in older Chinese adults remains limited.
Objective:
This study aimed to test whether adding a safety user interface (UI) bundle to a generative AI chat interface increases verification intention among older Chinese adults and to examine selected secondary outcomes, including reliance intention, trust calibration, perceived trustworthiness, comprehension, usability/readability, cognitive load, and a behavioral proxy of verification.
Methods:
We conducted a cross-sectional survey with an embedded randomized UI vignette experiment between May 22, 2025, and September 3, 2025. Chinese adults aged ≥60 years were recruited through community sites, outpatient clinic waiting areas, and WeChat (Tencent Holdings Ltd) groups, and randomized 1:1 to view screenshots of a baseline chat UI or a safety UI bundle containing generic source-label cues, and an uncertainty and verification nudge. Each participant completed 2 scenarios (service/travel decision and general well-being related to sleep/fatigue), followed by measures of verification intention (primary), reliance intention, trust calibration index, comprehension (0-8), perceived trustworthiness, usability/readability, cognitive load (0-10), manipulation checks, and a behavioral proxy (expanding optional "source information"). Analyses used intention-to-treat regression models with covariate adjustment.
Results:
Of 214 consenting respondents who started the survey, 200 were included in the analysis (100 per arm). The safety UI bundle increased verification intention (mean 4.72, SD 0.63 vs 4.41, SD 0.59 on a 7-point scale; adjusted β=0.293, 95% CI 0.128-0.457; P<.001). Reliance intention did not increase (mean 4.97, SD 0.54 vs 5.03, SD 0.58; adjusted β=-0.105, 95% CI -0.239 to 0.029; P=.13). Trust calibration improved (trust calibration index: mean -0.29, SD 1.43 vs 0.29, SD 1.43; adjusted β=-0.567, 95% CI -1.005 to -0.129; P=.01). Expansion of optional source information was numerically higher, although the adjusted CI included the null (42% vs 27%; adjusted odds ratio [OR]=1.76, 95% CI 0.95-3.27; P=.07). Comprehension remained high and similar across arms (mean 6.33, SD 1.14 vs 6.32, SD 1.08; adjusted β=-0.132, 95% CI -0.428 to 0.163; P=.38). Perceived trustworthiness was modestly lower in the Safety UI arm (mean 5.20, SD 0.61 vs 5.39, SD 0.66; adjusted β=-0.199, 95% CI -0.382 to -0.016; P=.03). Usability/readability was unchanged, and cognitive load did not increase. Manipulation checks indicated higher cue recognition in the Safety UI arm.
Conclusions:
In a randomized static-vignette survey of older Chinese adults, a brief safety UI bundle was associated with higher verification intention and a trust calibration index consistent with lower overreliance risk, without detectable reductions in comprehension or usability/readability. Because the intervention was tested as a bundle using screenshots and generic source labels, findings should be interpreted as evidence for a practical interface-level strategy rather than proof that any single cue caused the observed effects.