Related Experiment Video
Updated: Aug 12, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Prompt Framing and Evidence Requirement for AI-Generated Educational Responses in Dental Education: Experimental
Man Hung1,2,3, Corban Ward1, Owen Cohen1
1College of Dental Medicine, Roseman University of Health Sciences, 10894 S. River Front Parkway, South Jordan, UT, 84095, United States, 1 8018781270.
Background:
Large language models are increasingly used in health professions education; however, the role of prompt design in shaping their outputs remains poorly understood in clinical training contexts. In dentistry, where information presentation, perceived credibility, and procedural reasoning are important, the effects of instructional framing and evidence requirements on AI-generated educational responses are particularly relevant.
Objective:
This study examined whether instructional framing and evidence requirements were associated with rater-assessed perceived factuality, tone, stance orientation, citation behavior, safety notices, hedging, and response length in responses about cavity preparation generated using GPT-5 through its web-based interface.
Methods:
In a 2×2 factorial experiment, we manipulated instructional framing (patient-centered vs skill-centered) and evidence requirement (evidence-required vs no evidence) across 10 base prompt topics and 4 experimental conditions, yielding 40 outputs. Five trained raters independently coded each response for perceived factuality, confidence tone, stance orientation, hedging, citation presence, and safety notices. Response length was calculated programmatically. Interrater reliability was assessed using intraclass correlation coefficients and Fleiss κ. Consensus measures were analyzed using factorial analyses of variance and chi-square tests.
Results:
Evidence requirement was associated with greater citation presence (19/20, 95% vs 5/20, 25%; P<.001) and longer responses (P<.001). It was also associated with higher rater-assessed perceived factuality in exploratory analyses (P=.02). Hedging and confidence tone showed nonsignificant patterns. Instructional framing was associated with stance orientation (P=.006) but not with response length. No refusals occurred, and safety notices were infrequent across conditions. Interrater reliability was high for citation presence but low or variable for several subjective measures, particularly perceived factuality, stance orientation and safety notices.
Conclusions:
Prompt design was associated with differences in the presentation, structure, and orientation of large language model-generated educational responses in dentistry. Evidence requirements increased citation inclusion and response length, whereas instructional framing was associated with the stance emphasized in the response. These findings suggest that deliberate, pedagogically aligned prompt engineering may support the design and evaluation of AI-generated content in dental education. However, the effects of prompt wording on objectively verified accuracy, clinical safety, and learning outcomes require further investigation.