Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking Large Language Models for Scientific Writing: A Mixed‑Methods Evaluation of Chatgpt, Deepseek, and Human
Kulyash R Zhilisbayeva1, Nasrin Shokrpour2,3, Mohammad Mahdi Parvizi4,5
1Department of Languages, West Kazakhstan Marat Ospanov Medical University, Aktobe, Kazakhstan.
Introduction:
Modern large language models (LLMs) like ChatGPT (based on the GPT-4 architecture) and DeepSeek offer unprecedented capabilities for generating scientific text. However, their performance in replicating structured, high-quality scientific writing, especially compared to human-authored abstracts, remains insufficiently evaluated. To compare the abstract quality produced by human authors, ChatGPT/GPT4, and DeepSeek across: six evaluation criteria Clarity, Coherence, Conciseness, Accuracy, IMRaD Structure, and Language Quality, using blinded expert ratings and non-parametric statistical methods, specifically the Kruskal-Wallis test followed by pairwise Wilcoxon rank-sum tests with false discovery rate correction.
Methods:
We selected 23 medical and healthrelated research topics, each yielding three abstracts (human, ChatGPT, DeepSeek), for a total of 69 abstracts. Three raters scored each abstract. Kruskal-Wallis tests assessed group differences; Cliff's Delta (δ) was calculated as a nonparametric effect size for each comparison, suitable for ordinal data.
Results:
Across criteria, ChatGPT and DeepSeek significantly outperformed human authors in Clarity, Coherence, IMRaD Structure, and Language Quality. In contrast, Conciseness and Accuracy showed negligible effect sizes (|δ| <0.10), suggesting parity across all three sources.
Conclusions:
ChatGPT and DeepSeek achieved significantly higher scores in clarity, coherence, structure, and language quality, while showing comparable performance in conciseness and accuracy. These findings complement recent evaluations showing competitive medical and reasoning performance of DeepSeek models compared to proprietary LLMs. While shortform abstracts, expert oversight, and domain expertise remain critical, the results suggest that LLMs-particularly GPT4 and DeepSeek-can serve as effective tools in drafting scientific abstracts.
