Related Experiment Video
Updated: May 19, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Impact of prompt engineering on large language models for risk of bias assessment: a comparative study
Yu-Tao Xiong1, Yuan-Chen Wang1, Ya-Nan Sun2
1State Key Laboratory of Oral Diseases & National Centre for Stomatology & National Clinical Research Centre for Oral Diseases & Department of Oral and Maxillofacial Surgery, West China Hospital of Stomatology, Sichuan University, Chengdu, Sichuan, China.
Objectives:
To evaluate the performance of large language models (LLMs) in risk of bias assessment and to examine whether prompt engineering improves their accuracy and alignment with expert reasoning.
Methods:
We analysed 158 randomised controlled trials from 10 dental systematic reviews and their risk of bias assessments were reviewed and revised to serve as the reference standard. Two LLMs (DeepSeek-V3 and GPT-5) were evaluated under four prompting strategies, including direct command, command with reference, constrained output and formula-constrained output. The direct command served as the blank control group, simulating the approach commonly used by clinicians, whereas the other three groups employed different prompt engineering. The performance of LLMs across the seven domains of RoB-1 was evaluated using accuracy and agreement. The reasoning process of the LLMs was expressed in the form of syllogisms and its similarity to expert reasoning was assessed using MMD2.
Results:
LLMs showed limited capability in risk of bias assessment under the blank control condition, with mean accuracies of 0.72 for DeepSeek-V3 and 0.65 for GPT-5. With formula-constrained prompting, the performance of both LLMs improved significantly, and the overall accuracy increased to 0.85 for both DeepSeek-V3 and GPT-5 (both vs the blank control group, p<0.001). Agreement metrics showed a similar pattern, with higher agreement under formula-constrained prompting than under the other prompting strategies (p<0.001 for both models). In addition, the syllogistic output format provided a clear representation of the reasoning process underlying risk of bias assessment. Compared with constrained output, formula-constrained prompting also produced reasoning that was more closely aligned with the reference answers, as indicated by lower MMD² values (DeepSeek-V3: 0.0765 vs 0.1239; GPT-5: 0.0548 vs 0.1068).
Conclusion:
Prompt engineering substantially improved the performance of LLMs in risk of bias assessment. Although LLMs cannot currently replace human reviewers, they may serve as efficient and transparent tools to support this process.
Related Concept Videos
Bias in Epidemiological Studies
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Improving Translational Accuracy
Improving Translational Accuracy