Related Experiment Video
Updated: May 5, 2026

07:10
Application of Biochip Microfluidic Technology to Detect Serum Allergen-specific Immunoglobulin E sIgE
Published on: April 21, 2019
15.9K
ChatGPT Performance on 120 Interdisciplinary Allergology Questions-Systematic Evaluation With Clinical Error Impact
Sonja Mathes1, Sebastian Seurig2, Friederike Bluhme3
1Department for Dermatology and Allergology, School of Medicine, Technical University of Munich, Munich, Germany.
The Journal of Allergy and Clinical Immunology. in Practice
|March 29, 2025
Summary
ChatGPT shows good accuracy for allergy questions but has critical errors, especially in pediatrics. Expert medical advice remains crucial for patient safety in allergology.
Area of Science:
- Artificial Intelligence in Medicine
- Allergy and Immunology Research
- Clinical Decision Support Systems
Background:
- Patients increasingly use AI tools like ChatGPT for medical information, including allergy-related queries.
- Long wait times for allergology appointments drive patients to seek alternative information sources.
- ChatGPT, while accessible, may provide inaccurate or incomplete medical advice, posing risks.
Purpose of the Study:
- To systematically evaluate the performance of ChatGPT (3.5) in answering allergological questions from clinical practice.
- To develop and apply an Allergological Error Impact Assessment to rate the severity and consequences of AI-generated errors.
- To analyze ChatGPT's accuracy, completeness, perceived humanness, and readability in allergy-related contexts.
Main Methods:
- 120 multidisciplinary allergy questions (dermatology, pediatrics, pulmonology) were posed to ChatGPT.
- Responses were assessed for accuracy, completeness, perceived humanness, and readability (Flesch Reading Ease).
- Errors were categorized by severity (minor, major, critical), with critical errors undergoing impact analysis.
Main Results:
- ChatGPT achieved good accuracy (mean 4.1/5) but exhibited critical errors in 6% of responses (1 dermatology, 2 pediatrics, 3 pulmonology).
- Completeness and perceived humanness were lower for pediatric queries.
- A critical error concerning pediatric food allergens presented a potentially life-threatening risk.
Conclusions:
- ChatGPT's current reliability in allergology is imperfect, underscoring the necessity of expert medical consultation.
- AI tools require specialized tailoring for allergy use cases to enhance their utility in clinical settings.
- Further development could enable AI like ChatGPT to safely assist in routine allergy care.
Keywords:
AI in allergologyAdverse outcomes of ChatGPT health consultationAllergological counselingChatGPT health information accuracyChatGPT in allergologyError impact assessmentError ratingPatient perspectivesRoutine careMore Related Videos
Related Concept Videos
Random and Systematic Errors
11.2K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
11.2K
Systematic Error: Methodological and Sampling Errors
8.7K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
8.7K

