Related Experiment Videos
Comparing Artificial intelligence to physicians' competences in the domain of clinical reasoning: A systematic review
Mona Mlika1,2, Mohamed Majdi Zorgati3, Imen Ben Ismail2,4
1Department of Pathology, Center of Traumatology and Major Burns, Ben Arous, Tunis, Tunisia.
Background:
Clinical reasoning is a complex process that plays a crucial role in order to solve the patients' problems. Our aim was to compare AI competences to the physicians' competences in clinical reasoning.
Methods:
We performed a meta-analysis under the guidelines of the AMSTAR (version 2). All eligible articles were retrieved from Pubmed, Embase and Cochrane databases. The articles included were rated according to the MMERSQI scores. Binary diagnostic accuracy was used as the primary outcome. Continuous performance scores were considered as secondary outcome in order to avoid excluding studies using continuous scores instead of concordance rates. The Review Manager software 5.4 (free version) was used to conduct this meta-analysis. The OR with the 95% CI were calculated for studies reporting binary accuracy. The SMD with the 95% CI were calculated for studies using continuous scores. Q test and I2 statistics were carried out to explore the heterogeneity among studies. P value <0.1 for q test or I2 value >50% represented substantial between study heterogeneity. A random-effects model was used. Subgroup analyses were performed to explore the potential sources of heterogeneity if necessary. Publication bias was assessed using the funnel plot analysis.
Results:
Considering the studies using binary scores and comparing odds ratios, 1609 clinical vignettes were used to compare the different LLM to the human clinical reasoning. The combined OR reached 0.65 with 95% CI [0.38, 1.12]. No significant difference between both groups was observed (p=0.12). When considering studies using continuous scoring systems, the combined standard mean difference reached 0.08 with 95% CI [-0.19, 0.35]. No significant difference between both groups was observed. In order to explain the heterogeneity that was noticed, we performed a sub-group analysis taking into account the nature of the clinical cases (real-world or published), the LLM system used (Chat-GPT v4) and the expertise of the respondants (novices or experts). The heterogeneity was observed in all subgroups excluding the subgroup of the studies using binary scoring systems and comparing LLM's scores to novices' scores. The combined OR reached 0.62 with 95% CI [0.35, 1.1]. No significant difference was observed between both groups (p=0.1) and the heterogeneity I-square was evaluated to 0% and Tau2 to 0.00.
Conclusion:
Even if this meta-analysis showed the absence of difference between AI and human clinical reasoning, these reaults have to be taken with caution because of the important heterogeneity that wasn't resolved by subgroup analyses.
Related Concept Videos
Critical Thinking II
Patient-centered Care
Critical Thinking I
Reason and Intuition
Ethical Dilemmas I
Let us explore some examples to understand the potentially complex moral decisions nurses face.
Take the case of caring for minors, particularly in areas related to reproductive...
Current Trends in Nursing II