Related Experiment Video
Updated: May 8, 2026

A New Technique for Quantitative Analysis of Hair Loss in Mice Using Grayscale Analysis
Published on: March 9, 2015
Generative Artificial Intelligence Tools: Evaluating Ways to Automate Your SALT (GATEWAYS) Scoring of Alopecia Areata
Radhika Gupta1, Hojjat Salmasian2, Michelle Oboite1,3
1Perelman School of Medicine, University of Pennsylvania, Philadelphia, Pennsylvania, USA.
Background:
Alopecia areata (AA) is an autoimmune disease affecting hair follicles that results in nonscarring hair loss. AA impacts 0.1%-0.2% of the United States population, with pediatric patients accounting for 16.0%-27.7% of all cases. The Severity of Alopecia Tool (SALT), a method of quantifying scalp alopecia, helps guide clinical practice and determine response to therapies in clinical trials. Given the emerging role of image-based assessments of alopecia and growth of multimodal generative artificial intelligence (AI) in dermatology, we aimed to assess the "off-the-shelf" ability of a large-language model, GPT-4o, to automate the generation of image-based SALT scores.
Methods:
Chart review of patients with AA seen at the Children's Hospital of Philadelphia's Dermatology Clinic was conducted to identify 4-view images of patients' scalps and provider-derived SALT scores. One-hundred-and-four 4-view image sets were de-identified and provided to GPT-4o, which was prompted to generate SALT scores. Concordance between GPT-4o's and providers' scores was determined using intraclass correlation coefficients (ICC) and concordance correlation coefficients (CCC).
Results:
ICC and CCC between GPT-4o and in-person provider assessments were 0.815 and 0.866. ICC and CCC between GPT-4o and image-based provider assessments were 0.833 and 0.817. ICC and CCC between two providers were 0.950 and 0.948. These high levels of concordance were confirmed on Bland-Altman plots.
Conclusions:
SALT scoring for AA can be challenging due to provider subjectivity, changing sphericity and growth of patients' scalps, particularly among pediatric patients. Our data show the potential adjunct role that "off-the-shelf" generative AI tools may play in SALT scoring without any prior additional explicit training.

