Related Experiment Video
Updated: Jun 25, 2025

10:32
Development of a Virtual Reality Assessment of Everyday Living Skills
Published on: April 23, 2014
18.5K
Controlling Inputter Variability in Vignette Studies Assessing Web-Based Symptom Checkers: Evaluation of Current
András Meczner1,2, Nathan Cohen1, Aleem Qureshi1
1Healthily, London, United Kingdom.
JMIR Formative Research
|May 31, 2024
Summary
Standardizing vignette testing for symptom checkers (SCs) improves tester agreement but significant variability remains. Isolated performance metrics offer a more reliable assessment than composite accuracy alone.
Area of Science:
- Digital health
- Health informatics
- Medical AI
Background:
- Web-based symptom checkers (SCs) are rapidly growing without standardized quality assurance.
- Vignette studies are common for SC evaluation, but accuracy is a composite metric influenced by tester variability.
- Existing methods lack widely accepted criteria for assessing SC performance.
Purpose of the Study:
- To evaluate the impact of tester variability on SC outcome accuracy using clinical vignettes.
- To explore the feasibility of measuring isolated aspects of SC performance.
- To improve the reliability and generalizability of SC performance assessments.
Main Methods:
- Assessed Healthily's SC with 114 vignettes across three tester groups (free, partially free, restricted instructions).
- Calculated κ statistics for agreement on outcome condition and triage.
- Measured crude and adjusted accuracy against a gold standard, with adjusted accuracy refined by a review process.
- Assessed symptom comprehension feasibility using variations of 51 chief complaints across three SCs.
Main Results:
- Intertester agreement improved with increased standardization (free: 0.49-0.51; partially free: 0.66-0.66; restricted: 0.71-0.72).
- Restricted group accuracy averaged 50.6% (SD 5.35%), with adjusted accuracy at 56.1%.
- Symptom comprehension assessment was feasible, with scores ranging from 52.9% to 68%.
Conclusions:
- Standardizing vignette testing significantly enhances intertester agreement but does not eliminate tester-dependent variability.
- Composite accuracy measures in vignette studies are limited by tester variability and small sample sizes.
- Measuring isolated SC performance aspects, like symptom comprehension, provides more reliable assessments.
- An adjusted accuracy measure and isolated metrics are recommended for future SC performance evaluations.
Related Concept Videos
Variability: Analysis
140
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
140
Accuracy and Errors in Hypothesis Testing
196
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
196
Sensitivity, Specificity, and Predicted Value
284
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
284
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Random and Systematic Errors
10.9K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
10.9K
Systematic Error: Methodological and Sampling Errors
1.5K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
1.5K

