Related Experiment Videos
Benchmarking AI-Powered Translation of the EQ-5D-5L Patient-Reported Outcome Measure Using Automated Metrics:
Himanshu Vashisht1,2, Tomás Ward1,2, Willie Muehlhausen2,3
1School of Computing, Dublin City University, Collins Ave Ext, Whitehall, Dublin, Ireland, 353 899586131.
JMIR AI
|August 4, 2026
Summary
AI translation services show high semantic similarity for patient-reported outcome measures (PROMs) but require human review for conceptual equivalence. This supports AI in PROM workflows, though sensitive phrasing needs further research.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Clinical Research
Background:
- Patient-reported outcome measures (PROMs) are crucial for multinational clinical research.
- High-quality translation and linguistic validation of PROMs are resource-intensive.
- AI-powered translation offers potential acceleration but requires performance evaluation.
Purpose of the Study:
- To benchmark the quality and comparability of 4 AI translation services for the EQ-5D-5L.
- To evaluate AI translations against official, linguistically validated human translations (gold standard).
- To assess AI performance across 5 European languages: Danish, Dutch, French, German, and Spanish.
Main Methods:
- Translated EQ-5D-5L segments using Google Translate, GPT-4.1, Amazon Translate, and DeepL.
- Benchmarked AI outputs against human translations using BLEU, METEOR, COMET, and BLEURT metrics.
- Employed Friedman tests and post hoc analyses to identify statistically significant differences.
- Conducted descriptive and visual analyses for semantic similarity and localized deviations.
Main Results:
- Statistically significant performance differences between AI services were found in 55% of metric-language combinations.
- Most detected differences were stylistic rather than meaning-altering, identified by surface-overlap metrics.
- High semantic similarity was observed across services, with deviations concentrated in headers and abstract health concepts.
Conclusions:
- AI translation services demonstrate high semantic similarity for PROMs in high-resource European languages.
- AI can aid PROM translation workflows, but expert human review is essential for conceptual equivalence.
- Further research should focus on clinically sensitive phrasing and abstract health concepts.