Benchmarking Large Language Models for Scientific Writing: A MixedMethods Evaluation of Chatgpt, Deepseek, and Human

Kulyash R Zhilisbayeva1, Nasrin Shokrpour2,3, Mohammad Mahdi Parvizi4,5

  • 1Department of Languages, West Kazakhstan Marat Ospanov Medical University, Aktobe, Kazakhstan.

Summary

Large language models like ChatGPT and DeepSeek significantly improve scientific abstract quality in clarity, coherence, structure, and language compared to human authors. However, conciseness and accuracy remain comparable across all sources.

Related Concept Videos