Related Experiment Videos
Domain-Dependent Performance of Human Experts and AI Systems in Pediatric Menu Evaluation
Roxana Maria Martin-Hadmaș1, Rebeca Sovea1, Diana Pol1
1Department of Community Nutrition and Food Safety, "George Emil Palade" University of Medicine, Pharmacy, Science and Technology of Târgu Mures, Gheorghe Marinescu 38, 540139 Targu Mures, Romania.
Background/Objectives:
Artificial intelligence (AI) tools are increasingly used for dietary assessment, but their reliability in pediatric nutrition remains uncertain. This study compared AI systems and human evaluators in pediatric menu assessment against expert-defined reference ratings.
Methods:
This observational cross-sectional study used anonymized dietary and clinical information from three healthy pediatric cases. A multidisciplinary panel of pediatric nutrition specialists established standard evaluations. Assessments were completed by 84 AI evaluations, 116 nutrition specialists, 56 pediatric-focused physicians, and 88 physicians from other specialties. Outcomes included absolute error in total energy estimation, deviations in portion and qualitative ratings, standardized scores, and binary adequacy accuracy. Groups were compared using ANOVA or Kruskal-Wallis tests with post hoc analyses.
Results:
A total of 344 assessments were included. Energy-estimation error differed between groups (Kruskal-Wallis = 12.85, p = 0.0016), but this analysis was based on a limited and highly unbalanced subset of 77 assessments. AI had the largest mean absolute error (260.3 ± 194.3 kcal), followed by nutrition specialists (148.4 ± 81.5 kcal); physicians with other specialties had the lowest error (60.4 ± 53.3 kcal); these subgroup comparisons should be interpreted cautiously because of the small and unequal numbers of available energy estimates. AI showed greater portion-rating deviation (0.488; 95% CI: 0.38-0.60) than nutrition specialists (0.310; 95% CI: 0.20-0.42; p = 0.0019, q = 0.0059) and pediatric-focused physicians (0.286; 95% CI: 0.16-0.41; p = 0.0181, q = 0.019). Conversely, AI showed the lowest deviations for variety (0.37 ± 0.49) and processing (0.33 ± 0.47), compared with 1.15-1.36 and 1.22-1.57, respectively, among human groups (both p = 0.0001).
Conclusions:
AI may support structured pediatric menu screening for descriptive qualitative features; however, lower precision for energy and portion assessment supports, under these specific conditions, its use as an adjunct, for qualified nutrition professionals.