Related Experiment Video
Updated: Jun 10, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Performance of 3 large language models in detecting urinary formed elements
1Department of Clinical Laboratory Medicine, National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital & Shenzhen Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Shenzhen, China.
Large language models (LLMs) show promise for identifying urinary formed elements in unstained images. Google Gemini demonstrated the highest accuracy, though current models are not yet suitable for clinical diagnosis.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Laboratory Diagnostics
- Urine Analysis
Background:
- Large language models (LLMs) are increasingly integrated into medical applications, particularly for image recognition.
- Research on LLM capabilities for identifying unstained medical images, such as urinary formed elements, is limited.
Purpose of the Study:
- To evaluate and compare the image recognition performance of three LLMs (ChatGPT-4o, Deepseek Janus Pro7B, and Google Gemini) in detecting urinary formed elements.
- To assess the accuracy and consistency of these LLMs in analyzing urine morphology images.
Main Methods:
- A cross-sectional study analyzed 45 urine morphology images using a standardized prompt for ChatGPT-4o, Deepseek Janus Pro7B, and Google Gemini.
- Each image was evaluated three times by each LLM, with results assessed on a 5-point Likert scale.
- Statistical analysis included the Friedman test, Kendall W score, and Mann-Whitney U test to compare model performance.
Main Results:
- Google Gemini Advanced achieved the highest accuracy at 31%.
- ChatGPT-4o and Google Gemini demonstrated strong consistency (Kendall W scores of 0.797 and 0.812, respectively), outperforming Deepseek (W=0.663).
- Significant performance differences were observed, with Gemini excelling in identifying urinary casts and microorganisms.
Conclusions:
- While LLMs like ChatGPT-4o, Deepseek Janus Pro7B, and Google Gemini show potential for identifying urinary formed elements, their current diagnostic accuracy is insufficient for clinical use.
- Further advancements in training data and model optimization are necessary to enhance their clinical applicability in urine morphology analysis.
- Future research should focus on improving LLM performance for broader utility in medical laboratories.
Related Concept Videos
Filtration and Urine Formation
Urine Studies I: Urinalysis
Urinary Tract Infection III: Diagnostic Studies and Interprofessional Care
Formation of Dilute Urine
Filtrate Osmolarity in the PCT
Initially, as the filtrate passes through the proximal convoluted tubule (PCT), its...
Urinary Bladder
In males, the bladder is situated in front of the rectum, while in females, it is positioned anterior to the vagina and uterus. The bladder floor contains an inverted triangular area called the trigone, defined by the two ureteric...
Urodynamic Studies: Uroflowmetry
