Related Experiment Video
Updated: Aug 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking Large Language Models for Scientific Writing: A Mixed‑Methods Evaluation of Chatgpt, Deepseek, and Human
Kulyash R Zhilisbayeva1, Nasrin Shokrpour2,3, Mohammad Mahdi Parvizi4,5
1Department of Languages, West Kazakhstan Marat Ospanov Medical University, Aktobe, Kazakhstan.
Summary
Large language models like ChatGPT and DeepSeek significantly improve scientific abstract quality in clarity, coherence, structure, and language compared to human authors. However, conciseness and accuracy remain comparable across all sources.
Area of Science:
- Artificial Intelligence
- Scientific Communication
- Medical Research
Background:
- Large language models (LLMs) demonstrate advanced text generation capabilities.
- The efficacy of LLMs in producing high-quality scientific abstracts requires thorough evaluation.
- Comparing LLM-generated abstracts to human-authored ones is crucial for understanding their utility.
Purpose of the Study:
- To evaluate and compare the quality of scientific abstracts generated by humans, ChatGPT (GPT-4), and DeepSeek.
- To assess performance across six key criteria: Clarity, Coherence, Conciseness, Accuracy, IMRaD Structure, and Language Quality.
- To utilize blinded expert ratings and non-parametric statistical methods for objective comparison.
Main Methods:
- A total of 69 abstracts were collected across 23 medical and health-related research topics.
- Each topic included one abstract authored by a human, one by ChatGPT, and one by DeepSeek.
- Three expert raters scored each abstract, with Kruskal-Wallis tests and Cliff's Delta used for statistical analysis.
Main Results:
- ChatGPT and DeepSeek significantly outperformed human authors in Clarity, Coherence, IMRaD Structure, and Language Quality.
- No significant differences were observed in Conciseness and Accuracy, with negligible effect sizes (|δ| < 0.10).
- LLMs demonstrated parity with human authors in specific quality metrics.
Conclusions:
- LLMs, specifically ChatGPT (GPT-4) and DeepSeek, show potential as effective tools for drafting scientific abstracts.
- While LLMs excel in clarity, coherence, structure, and language, human oversight remains essential.
- The findings align with recent research highlighting the competitive performance of models like DeepSeek in medical and reasoning tasks.
