Related Experiment Video
Updated: Mar 1, 2026

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
What is the evidence for the performance of generic preference-based measures? A systematic overview of reviews
Aureliano Paolo Finch1, John Edward Brazier2, Clara Mukuria2
1Health Economics and Decision Science, School of Health and Related Research, University of Sheffield, West Court, 1 Mappin Street, Sheffield, S1 4DT, UK. APFinch1@sheffield.ac.uk.
Objective:
To assess the evidence on the validity and responsiveness of five commonly used preference-based instruments, the EQ-5D, SF-6D, HUI3, 15D and AQoL, by undertaking a review of reviews.
Methods:
Four databases were investigated using a strategy refined through a highly sensitive filter for systematic reviews. References were screened and a search for grey literature was performed. Identified citations were scrutinized against pre-defined eligibility criteria and data were extracted using a customized extraction template. Evidence on known group validity, convergent validity and responsiveness was extracted and reviewed by narrative synthesis. Quality of the included reviews was assessed using a modified version of the AMSTAR checklist.
Results:
Thirty reviews were included, sixteen of which were of excellent or good quality. The body of evidence, covering more than 180 studies, was heavily skewed towards EQ-5D, with significantly fewer studies investigating HUI3 and SF-6D, and very few the 15D and AQoL. There was also lack of head-to-head comparisons between GPBMs and the tests reported by the reviews were often weak. Where there was evidence, EQ-5D, SF-6D, HUI3, 15D and AQoL seemed generally valid and responsive instruments, although not for all conditions. Evidence was not consistently reported across reviews.
Conclusions:
Although generally valid, EQ-5D, SF-6D and HUI3 suffer from some problems and perform inconsistently in some populations. The lack of head-to-head comparisons and the poor reporting impedes the comparative assessment of the performance of GPBMs. This highlights the need for large comparative studies designed to test instruments' performance.
More Related Videos
13:20Online Repetitive Transcranial Magnetic Stimulation of Dorsomedial and Dorsolateral Prefrontal Cortex in Cognition Decision Making, and Cognitive Dissonance
Published on: December 5, 2025
08:27Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
Published on: September 27, 2019
Related Concept Videos
Stereotypes, Prejudice, and Discrimination
Stereotype Content Model
Social Proof
Bioequivalence: Overview
Bioequivalence Data: Statistical Interpretation
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...