Related Experiment Video
Updated: Jun 21, 2026

Development of a Virtual Reality Assessment of Everyday Living Skills
Published on: April 23, 2014
Inter-rater reliability and clinical utility of the modified ferriman-gallwey scale in routine practice
Irena Chongaroonngamsang1, Phawat Matemanosak1, Chitkasaem Suwanrath1
1Department of Obstetrics and Gynaecology, Faculty of Medicine, Prince of Songkla University, Songkhla, Thailand.
Objectives:
To evaluate the reliability and clinical utility of the modified Ferriman-Gallwey (mFG) score in routine practice when applied by clinicians with differing experience and by patients for self-assessment.
Methods:
In this cross-sectional reliability study, 43 reproductive-aged women were assessed for hirsutism. Three clinicians (reproductive endocrinologist, general gynecologist, and resident) independently scored participants using the mFG scale, and patients self-graded their hirsutism. Inter-rater reliability was examined using intraclass correlation coefficients (ICC) and weighted kappa statistics by body area. Clinical agreement was assessed with Bland-Altman limits of agreement (LOA).
Results:
Inter-rater reliability was good (ICC = 0.649), higher between experienced clinicians (ICC = 0.723) and lower when the resident was included (ICC = 0.588). Weighted kappa ranged from 0.118 to 0.832 (highest for chin, lowest for upper back). Bland-Altman analysis showed substantial disagreement: experienced pair bias -0.74 (LOA -6.94 to + 5.46) and less-experienced pair bias + 0.44 (LOA -7.22 to + 8.11). Patient self-scoring showed poor agreement with clinician scoring (ICC = 0.20).
Conclusions:
Although clinician scoring demonstrated good statistical reliability, inter-rater differences remained clinically meaningful, with LOA exceeding diagnostic thresholds for hirsutism. Greater experience improved agreement but did not eliminate variability. Patient self-scoring showed low agreement with experts and tended to overestimate severity; however, it may function as a high-sensitivity screening adjunct rather than a substitute for clinician assessment.
