Related Experiment Videos
[Development and psychometric validation of the User Experience Scale for Artificial Intelligence Question-Answering
Ruixi Liu1, Bifen Kuang2, Yifan Yang3
1Department of Stomatology, Second Xiangya Hospital, Central South University, Changsha 410011, China. Liurx211@163.com.
Objectives:
Oral health is an essential component of overall health; however, specific assessment tools for first-visit patients in dental outpatient waiting scenarios are currently lacking. This study aims to systematically develop and validate the User Experience Scale for Artificial Intelligence Question-Answering Systems (UES-AIQA) for Chinese first-visit dental patients, providing a standardized instrument for evaluating the service quality of artificial intelligence (AI)-based services in dental outpatient settings.
Methods:
An initial item pool was developed through literature analysis and expert group discussions. Item selection was conducted through 2 rounds of Delphi expert consultation. A total of 305 first-visit dental patients in the Department of Stomatology were recruited using convenience sampling. The first sample of 233 participants was used for exploratory factor analysis (EFA), and the second sample of 72 participants was used for confirmatory factor analysis (CFA). Psychometric properties of the scale were evaluated through item analysis, factor analysis, and reliability analysis.
Results:
The final scale consisted of 10 items and demonstrated a unidimensional structure. The item-level content validity index (I-CVI) ranged from 0.900 to 1.000, and the average scale-level content validity index (S-CVI/Ave) was 0.980. Parallel analysis and the Kaiser criterion both supported the extraction of one common factor. The cumulative variance contribution rate of the single factor was 83.593%, and standardized factor loadings of all items ranged from 0.879 to 0.937. CFA showed that the ratio of chi-square to degrees of freedom (χ2/df) was 3.190, the comparative fit index (CFI) was 0.929, the Tucker-Lewis index (TLI) was 0.900, and the standardized root mean square residual (SRMR) was 0.035, with the major fit indices reaching acceptable levels. The root mean square error of approximation (RMSEA) was 0.174. Standardized factor loadings in the CFA model ranged from 0.845 to 0.928. The Cronbach's α coefficient of the total scale was 0.981, McDonald's ω coefficient was 0.983, Spearman-Brown split-half reliability coefficient was 0.970, and the test-retest intra-class correlation coefficient (ICC) was 0.830.
Conclusions:
The UES-AIQA demonstrated satisfactory content validity and reliability according to psychometric standards, while structural validity indicators reached acceptable levels. The scale may serve as a preliminary standardized tool for evaluating user experiences with AI question-answering systems among first-visit dental patients.