Related Experiment Videos
Artificial Intelligence in Medical Writing: Subtle Errors and Their Complex Consequences
Shigeki Matsubara1,2,3, Daisuke Matsubara4,5, Kazuhiko Kotani4
1Department of Obstetrics and Gynecology, Jichi Medical University, Tochigi, Japan.
Abstract:
Although artificial intelligence (AI; such as ChatGPT) can improve manuscript readability, AI-related false descriptions (so-called hallucinations) and incorrect reference retrieval have been repeatedly reported. We tested (1) whether ChatGPT, when provided with medically correct inputs, generates a manuscript containing medically incorrect statements-particularly whether it misunderstands or overlooks "subtle but important" medical issues-and (2) whether ChatGPT cites inappropriate references. We input bullet points on placenta percreta, tasked ChatGPT-5 with generating a mini-review, and asked it to confirm whether the output was medically correct. We introduced a small, deliberate trap. In percreta, a recent conceptual change has gained increased attention: this condition is considered to result from uterine abnormality rather than abnormal placental invasion. This etiopathological shift could be misinterpreted as implying "weaker adherence" and therefore "less difficult surgery," leading to the notion that percreta could be managed at secondary-level institutions. The ChatGPT-generated manuscript cited appropriate references and was almost medically correct, except for one critical issue: it stated that percreta could be managed in secondary-level institutions, which is incorrect and potentially dangerous. During the "confirmation" stage, ChatGPT raised a caution regarding this issue, but not in a definitive manner. An additional experiment was conducted on hypercholesterolemia. ChatGPT again failed to address an important issue, familial hypercholesterolemia, which requires a management strategy different from that for non-familial hypercholesterolemia. Overall, ChatGPT generated a linguistically appealing manuscript with largely correct context, but it produced incorrect and potentially dangerous statements in clinically critical areas. When incorrect statements are subtle rather than obvious, authors, journals, and readers may fail to recognize them. This is paradoxical: advances in AI may reduce "apparent" errors while generating less recognizable ones. When evaluating AI-assisted manuscripts, careful review by individuals with deep domain knowledge is mandatory.
Related Concept Videos
Non-equilibrium in the Cell
Errors occurring during blood pressure monitoring
Several factors...
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Pharmaceutical Poisoning: Potential Scenarios
RNA Editing
Mismatch Repair