Related Experiment Video
Updated: May 11, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Performance and Safety of Large Language Models in Generating Type 2 Diabetes Mellitus Management
Agnibho Mondal1, Arindam Naskar2, Bhaskar Roy Choudhury3
1Department of Infectious Diseases and Advanced Microbiology, School of Tropical Medicine, Kolkata, IND.
Large language models like GPT-4 can reduce unnecessary prescriptions for type 2 diabetes patients but lack physician completeness. AI shows promise as a supplementary tool with human oversight for safe clinical use.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Clinical Decision Support
Background:
- Large language models (LLMs) offer potential in healthcare, from scientific writing to personalized medicine.
- Clinical application of LLMs requires rigorous evaluation for accuracy, ethics, and bias.
- Assessing LLM utility and safety against medical standards is crucial.
Purpose of the Study:
- To compare the clinical management plans generated by GPT-4 against those of physicians for type 2 diabetes mellitus patients.
- To evaluate the completeness, necessity, and dosage accuracy of AI-generated versus physician-generated management plans.
- To assess the safety of GPT-4's clinical management plans.
Main Methods:
- Comparative analysis of anonymized patient records from West Bengal, India.
- GPT-4 and three physicians generated management plans for 50 type 2 diabetes mellitus patients.
- Plans were evaluated against American Diabetes Association guidelines, quantifying completeness, necessity, and dosage accuracy.
Main Results:
- Physicians' plans had fewer missing medications (p=0.008), while GPT-4 had fewer unnecessary medications (p=0.003).
- No significant difference in drug dosage accuracy (p=0.975) or overall error scores (p=0.301) between GPT-4 and physicians.
- GPT-4 generated plans had safety issues in 16% of cases.
Conclusions:
- GPT-4 effectively reduces unnecessary prescriptions but does not match physician completeness in diabetes management.
- LLMs show potential as supplementary tools in clinical settings.
- Enhanced algorithms and continuous human oversight are essential for safe and effective AI implementation in healthcare.
Related Concept Videos
Diabetes Mellitus: Type 2 and Gestational
Diabetes: Management and Pharmacotherapy
Insulin remains the cornerstone of treatment for most patients with type 1 and many...
Diabetes Mellitus: Overview and Type I Subtype
Type 1 diabetes is an autoimmune disease in which the immune system mistakenly attacks and destroys the insulin-producing beta cells in the pancreas. As a result, the body is unable to produce sufficient insulin, and individuals with...
Carbohydrate Metabolism
Starch accounts for approximately 60% of the carbohydrates consumed by humans. Since amylase enzymes cannot function in the stomach's acidic environment, starch can only be digested in the mouth and small intestine. Simple sugars are found naturally in milk and fruits in...
Oral Hypoglycemic Agents: Biguanides and Glitazones
Diabetes: Symptoms, Diagnosis, and Complications

