Related Experiment Video
Updated: May 13, 2026

08:12
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
A comparison model for measuring individual agreement.
Lawrence Lin1, A S Hedayat, Yuqing Tang
1Baxter Healthcare Co., Round Lake, Illinois, USA.
Journal of Biopharmaceutical Statistics
|February 27, 2013
Summary
This study introduces a new model for assessing rater agreement with multiple readings. It proposes two indices, TIR and IIR, to compare total-rater and intrarater agreement for reliable scientific data.
Area of Science:
- Biostatistics
- Data Analysis
- Scientific Measurement
Background:
- Assessing inter-rater and intra-rater reliability is crucial in scientific studies.
- Existing methods may not adequately compare agreement across multiple raters and replicated readings.
- A flexible model is needed to evaluate individual agreement comprehensively.
Purpose of the Study:
- To propose a general comparison model for assessing individual agreement among multiple raters with replicated readings.
- To introduce two novel comparative agreement indices: Total-Intra Ratio (TIR) and Intra-Intra Ratio (IIR).
- To provide a framework for exploring total-rater versus intrarater agreement and comparing intrarater agreement among raters.
Main Methods:
- Development of a general comparison model based on mean squared deviations (MSDs).
- Introduction of TIR for noninferiority assessment of inter-rater agreement relative to intrarater agreement.
- Introduction of IIR for classical assessment of rater precision.
- Utilizing Generalized Estimating Equations (GEE) methodology for estimation and statistical inference.
Main Results:
- The proposed model allows flexible comparison of total-rater and intrarater agreement.
- TIR provides a noninferiority assessment applicable with or without a reference standard.
- IIR enables direct comparison of precision among selected raters.
- The FDA's bioequivalence method is shown as a special case of the TIR approach.
Conclusions:
- The proposed model and indices offer a robust framework for evaluating individual rater agreement in studies with multiple raters and readings.
- TIR and IIR provide valuable metrics for ensuring data reliability and understanding rater performance.
- The GEE methodology ensures sound statistical inference for the proposed indices.
Related Concept Videos
Kendall's Coefficient of Concordance
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects or...
Statistical Analysis: Overview
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Testing a Claim about Standard Deviation
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Multiple Comparison Tests
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Calibration Curves: Correlation Coefficient
In a linear calibration curve, there is a value called the calibration coefficient, denoted by 'r,' which measures the strength and the direction of association between two variables. The correlation coefficient value ranges from −1 to +1. A value of +1 indicates a perfect positive linear correlation, −1 denotes a perfect negative correlation, and 0 implies no correlation between the two variables. A positive correlation value establishes that as one variable increases, the other increases, and...
One-Way ANOVA: Equal Sample Sizes
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...

