Related Experiment Videos
AI Chatbot Answers for Drug Dosing Adjustments According to Renal Function in Geriatric Patients Using the New
Celine Barbonus1,2, Ralf Sultzer3, Thilo Bertsche1,2
1Department of Clinical Pharmacy, Institute of Pharmacy, Faculty of Medicine, Leipzig University, Brüderstraße 32, Leipzig, Saxony, 04103, Germany, 49 3419711800.
This study evaluated how well four popular AI chatbots could recommend correct drug dosage adjustments for elderly patients with kidney issues. Researchers created a new scoring system to measure the accuracy and safety of these AI responses. The results showed that the chatbots struggled more as patient health became more complex, and their performance was not reliable enough for safe use in hospitals.
Area of Science:
- Geriatric pharmacology and AI Quality Output Score assessment
- Clinical decision support systems within digital health
Background:
Preventable medication errors frequently harm elderly individuals, particularly when kidney filtration rates are reduced. Clinicians often struggle to manage complex drug regimens for these vulnerable populations. Artificial intelligence chatbots are increasingly proposed as potential assistants for calculating precise medication adjustments. No prior work had resolved whether these digital tools maintain sufficient accuracy for clinical decision support. That uncertainty drove the development of a standardized assessment framework for evaluating automated medical advice. Existing metrics lacked the specificity required to judge complex pharmacological recommendations generated by large language models. This gap motivated a rigorous investigation into the reliability of chatbot-generated dosing guidance. Researchers needed a robust method to quantify the safety and precision of these machine-generated outputs.
Purpose Of The Study:
The aim of this study was to evaluate the reliability of artificial intelligence chatbots in providing drug dosing adjustments for geriatric patients. Researchers sought to determine if these models could safely handle complex medication regimens for individuals with impaired kidney function. The investigation addressed the critical need for scientific validation of automated medical advice tools. No prior work had established a standardized scoring system to quantify the quality of such digital outputs. This gap motivated the development of the AQUOS metric to assess accuracy and safety. The study examined whether performance depended on patient health status, language, or medication complexity. Investigators also aimed to identify the potential for patient harm resulting from inaccurate machine-generated recommendations. The project ultimately sought to clarify whether these technologies are ready for deployment in high-risk clinical environments.
Main Methods:
Review approach involved a cross-sectional analysis of four distinct large language models. Investigators utilized a standardized query format to test drug dosing recommendations for one hundred elderly patients. The team prompted these models in both English and German to evaluate language-based performance variations. Researchers assessed each model at two separate time points to determine the consistency of the generated advice. The study employed the newly established AQUOS metric to grade the precision of every response. Experts also applied the World Health Organization safety framework to identify potential risks within the provided guidance. This systematic evaluation covered sixteen hundred total interactions to ensure statistical power. The design focused on comparing performance metrics across varying levels of patient health complexity.
Main Results:
Key findings from the literature indicate that chatbot performance scores ranged from -19.0% to 95.2% across the tested models. The primary observation was that accuracy significantly declined as renal function decreased, with ChatGPT showing a correlation of -0.215. Medication complexity also negatively impacted results, specifically for Scite, which demonstrated a correlation of -0.239. The researchers noted that potential harm increased when patient statuses involved both lower kidney function and higher drug counts. English-language prompts yielded scores up to 4.8% higher than those provided in German. The generated advice proved highly reproducible when tested at two independent time intervals. Despite these findings, the overall quality remained insufficient for safe application in clinical settings. The data confirm that these tools struggle to handle the intricacies of geriatric drug management.
Conclusions:
The researchers propose that current chatbot performance remains inadequate for deployment in high-risk medical environments. Synthesis and implications suggest that automated tools fail to provide reliable guidance for patients with significant renal impairment. The study demonstrates that increasing clinical complexity consistently degrades the quality of machine-generated pharmacological advice. Authors emphasize that even the most accurate models do not meet the standards required for safe patient management. These findings highlight a persistent vulnerability when relying on artificial intelligence for critical drug dosing decisions. The data indicate that language differences marginally affect output quality, though this does not overcome broader reliability issues. Experts conclude that human oversight remains mandatory for all medication adjustments in geriatric care settings. Future efforts must address these performance limitations before such technology can be safely integrated into clinical workflows.
Frequently Asked Questions
The researchers propose that chatbot accuracy, measured by the AQUOS metric, significantly decreases as patient kidney function declines. Specifically, ChatGPT showed a negative correlation coefficient of -0.215, indicating that lower renal filtration leads to less reliable dosing recommendations.
The authors developed the AI Quality Output Score (AQUOS) to quantify performance. This tool evaluates responses on a scale from 0% to 100%, allowing for standardized comparisons across different models like ChatGPT, Copilot, Gemini, and Scite.
The researchers utilized a standardized prompt approach to query four distinct chatbots. This method was necessary to ensure that the input data remained consistent, allowing for a fair comparison of how each model processed complex geriatric medication scenarios.
The study analyzed 1600 individual responses. This large dataset allowed the team to assess reproducibility, language influence, and the potential for patient harm across various clinical scenarios involving polymedication.
The team measured potential harm using the World Health Organization's conceptual framework for patient safety. They observed that the risk of harmful advice increased when managing patients with both lower kidney function and higher medication complexity.
The authors state that even the highest scores achieved by these models are insufficient for clinical use. They argue that the current technology poses too great a risk to be deployed in high-stakes healthcare settings.
Related Concept Videos
Pharmacokinetics in Geriatric Patients: Effect of Age on Drug Excretion
Drug Dosing: Geriatric Patients
Drug Dosing in Renal Diseases: Dose Adjustments Based on Drug Clearance and Elimination Rate Constant
Drug Dosing in Renal Diseases: Estimation of Glomerular Filtration Rate Based on Serum Creatinine Concentration
Renal Failure: Dose Adjustments
Reduced renal clearance and elimination rate are common outcomes of renal impairment. These alterations lead to a prolonged elimination half-life and an altered apparent volume of distribution for drugs. As a result, dosage adjustments are typically necessary to maintain optimal drug levels in the body.
However, dosage adjustments...
Drug Dosing in Renal Diseases: Measurement of Glomerular Filtration Rate