A statistical simulation study finds discordance between WHO criteria and RECIST guideline

Madhu Mazumdar1, Alex Smith, Lawrence H Schwartz

  • 1Department of Epidemiology and Biostatistics, Memorial Sloan-Kettering Cancer Center, 307 E. 63rd St., 3rd Floor, New York, NY 10021, USA. mazumdam@mskcc.org

Abstract

Insights

The Response Evaluation Criteria In Solid Tumor (RECIST) guideline often leads to different cancer response classifications compared to the World Health Organization (WHO) standard. This discrepancy, particularly in progressive disease assessment, necessitates careful interpretation when evaluating new cancer therapies.

Area of Science:

  • Oncology
  • Clinical Trial Methodology
  • Biostatistics

Background:

  • Tumor shrinkage is a key endpoint for evaluating anticancer agents.
  • The World Health Organization (WHO) and Response Evaluation Criteria In Solid Tumor (RECIST) guidelines offer different methods for measuring tumor response.
  • RECIST, focusing on maximal diameter, was proposed to better correlate with tumor cell kill than WHO's product of diameters.

Purpose of the Study:

  • To statistically evaluate the concordance between WHO and RECIST criteria for tumor response assessment.
  • To determine the impact of tumor shape changes on response categorization differences.
  • To assess the clinical implications of discrepancies in response classification.

Main Methods:

  • Statistical simulation using tumor measurements and response data from 130 cancer patients.
  • Systematic assessment of concordance using Kappa coefficient and percentage disagreement.
  • Analysis across varying percentages of elliptical tumors and changes in tumor shape.

Main Results:

  • Overall disagreement between WHO and RECIST ranged from 14-20%.
  • Significant discrepancies were observed in progressive disease (32-35%) and partial response (8-16%) categories.
  • Tumor shape change magnitude significantly impacted concordance, while baseline ellipticity had minimal effect.

Conclusions:

  • RECIST and WHO criteria frequently result in different patient response categorizations, particularly for progressive disease.
  • These differences can complicate comparisons between new experimental therapies and historical conventional agents.
  • Increased stable disease (SD) under RECIST requires careful interpretation to differentiate criterion effects from actual drug-induced stabilization.

Related Concept Videos

Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches01:23

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches

Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Confounding in Epidemiological Studies01:27

Confounding in Epidemiological Studies

Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This phenomenon...
Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
Cochran's Q Test01:17

Cochran's Q Test

Cochran's Q Test is a nonparametric statistical test used to determine if there are potential differences in the outcomes of three or more related groups on a binary (yes/no) or dichotomous outcome. It is essentially an extension of the McNemar Test, which is limited to two related samples - Cochran's Q test can handle three or more related samples, making it more versatile in scenarios where subjects are measured under multiple conditions. The test statistic follows a Chi-Square distribution,...