Related Experiment Video
Updated: Jun 18, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
534
Comparison of Medical Research Abstracts Written by Surgical Trainees and Senior Surgeons or Generated by Large
Alexis M Holland1, William R Lorenz1, Jack C Cavanagh2
1Division of Gastrointestinal and Minimally Invasive Surgery, Department of Surgery, Atrium Health Carolinas Medical Center, Charlotte, North Carolina.
JAMA Network Open
|August 2, 2024
Summary
Artificial intelligence (AI) chatbots can generate medical research abstracts indistinguishable from human-written ones. While one AI model graded abstracts similarly to experts, another was less stringent, offering insights for AI implementation in research.
Area of Science:
- Medical research methodology
- Artificial intelligence in scientific writing
- Natural language processing applications
Background:
- Artificial intelligence (AI), particularly large language models like ChatGPT, is increasingly used in academia.
- However, its application and efficacy in generating and evaluating medical research abstracts remain under-explored.
- This study addresses the gap by assessing AI's capability in medical abstract creation and grading.
Purpose of the Study:
- To evaluate the ability of AI chatbots (ChatGPT versions 3.5 and 4.0) to generate medical research abstracts.
- To assess the performance of these AI models in grading medical research abstracts.
- To compare AI-generated abstracts and their grading against those produced by human researchers.
Main Methods:
- A cross-sectional study involving ChatGPT versions 3.5 (chatbot 1) and 4.0 (chatbot 2).
- AI models generated 10 abstracts using provided data and literature, alongside 10 human-written abstracts by surgical trainees and a senior author.
- Abstracts were evaluated and ranked by blinded surgeon-reviewers and the AI models themselves on 10- and 20-point scales.
Main Results:
- Surgeon-reviewers could not differentiate between AI-generated and human-written abstracts.
- No significant differences were observed in median scores or ranks between abstracts generated by residents, senior authors, or chatbot 1.
- Chatbot 2 graded abstracts more favorably than surgeon-reviewers and chatbot 1, indicating potential differences in AI grading stringency.
Conclusions:
- Trained AI chatbots can produce medical research abstracts comparable in quality to those written by human researchers.
- Chatbot 1's grading aligned with expert reviewers, while Chatbot 2 exhibited a more lenient grading approach.
- These findings suggest AI's potential utility for surgeon-scientists in medical research, with considerations for model-specific performance.

