Related Experiment Video
Updated: Jan 14, 2026

Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
Assessing the ability of ChatGPT 4.0 in generating check-up reports
Yikai Chen1, Yuxin Liu2, Yuanchang Huang1
1Department of Gastroenterological Surgery, The First Affiliated Hospital of Shantou University Medical College, Shantou, China.
ChatGPT 4.0 shows promise in generating accurate health reports for check-ups, excelling in guideline adherence and diagnosis but needing improvement in prioritizing high-risk items and providing comprehensive suggestions for better patient care.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Natural Language Processing
Background:
- Generative language models like ChatGPT (Chat Generative Pre-trained Transformer) are increasingly applied in clinical settings.
- Health check-ups are a common method for comprehensive personal health assessment.
- This study evaluates ChatGPT 4.0's utility in generating personalized health reports from check-ups.
Purpose of the Study:
- To assess the accuracy and personalization of health reports generated by ChatGPT 4.0.
- To evaluate ChatGPT 4.0's performance across different case complexities and languages.
- To determine the potential of ChatGPT 4.0 in enhancing clinical check-up efficiency and quality.
Main Methods:
- 89 ChatGPT 4.0-generated health check-up reports were analyzed.
- Reports were translated into English by ChatGPT 4.0 and graded by three doctors on six criteria (guidelines, diagnosis, order, system, consistency, suggestion).
- Case complexity (LOW, MEDIUM, HIGH) and language (English, Chinese) were factors in the analysis, using Wilcoxon rank sum and Kruskal-Wallis tests.
Main Results:
- ChatGPT 4.0 demonstrated strengths in guideline adherence, diagnostic accuracy, systematic presentation, and consistency.
- The model struggled with prioritizing high-risk items and providing comprehensive suggestions; some reports had data inconsistencies or were incorrect.
- Performance varied by complexity, with English reports showing differences across levels and Chinese reports showing distinct performance across all categories; no significant language advantage was found.
Conclusions:
- ChatGPT 4.0 can serve as an assistant for examiners, particularly for simpler tasks within health check-up reports.
- The model has the potential to improve medical efficiency and the quality of clinical check-up services.
- Further development is needed to enhance its ability to handle complex cases and provide more thorough recommendations.
Related Concept Videos
Methods of Documentation III: PIE
Guidelines and Strategies for Safe Computer Charting
Maintain Confidentiality and Security:
Data Collection III
The principles to begin the physical assessment include conducting a comprehensive or problem-related history in a quiet, well-lit room, emphasizing privacy and comfort for the...

