Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Friedman Two-way Analysis of Variance by Ranks01:21

Friedman Two-way Analysis of Variance by Ranks

267
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
267
Cochran's Q Test01:17

Cochran's Q Test

461
Cochran's Q Test is a nonparametric statistical test used to determine if there are potential differences in the outcomes of three or more related groups on a binary (yes/no) or dichotomous outcome. It is essentially an extension of the McNemar Test, which is limited to two related samples - Cochran's Q test can handle three or more related samples, making it more versatile in scenarios where subjects are measured under multiple conditions. The test statistic follows a Chi-Square...
461
Bonferroni Test01:10

Bonferroni Test

2.8K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.8K
Sign Test for Matched Pairs01:17

Sign Test for Matched Pairs

183
The sign test for matched pairs offers a robust method for comparing two paired samples, often for the effects of an intervention in one of them. This method is very useful in situations where the underlying distribution of the data is unknown. The test compares two related samples—often pre- and post-treatment measurements on the same subjects—to determine if there are significant differences in their median values.
To conduct the sign test, we first calculate the differences in...
183
Factorial Design02:01

Factorial Design

13.1K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.1K
One-Way ANOVA01:18

One-Way ANOVA

8.1K
One-way ANOVA analyzes more than three samples categorized by one factor. For example, it can compare the average mileage of sports bikes. Here, the data is categorized by one factor - the company. However, one-way ANOVA cannot be used to simultaneously compare the sample mean of three or more samples categorized by two factors. An example of two factors would be sports bikes from different companies driven in different terrains, such as a desert or snowy landscape. Here, two-way ANOVA is used...
8.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Neuromuscular Blockade Use and Monitoring Practices Reported by Anesthesiology Providers Practicing in Ambulatory Surgery Settings.

Anesthesia and analgesia·2026
Same author

Assessing a biomarker platform for stratifying indeterminate pulmonary nodules, protocol for the SPOT IT randomized trial.

Trials·2026
Same author

Postoperative residual neuromuscular blockade and hypoventilation after rocuronium and sugammadex or neostigmine for kidney transplantation: A randomized clinical trial.

Journal of clinical anesthesia·2026
Same author

Paediatric difficult intubations and the impact of timing of weekend versus weekday cases. Response to Br J Anaesth 2026; 136: 365-6.

British journal of anaesthesia·2026
Same author

Remote guidance and focused point-of-care ultrasound (POCUS) training for cardiopulmonary instability assessment by novice sonographers.

Resuscitation plus·2026
Same author

Randomized Controlled Trial of Upper Esophageal Sphincter Assist Device in Laryngopharyngeal Reflux Disease.

Clinical gastroenterology and hepatology : the official clinical practice journal of the American Gastroenterological Association·2025

Related Experiment Video

Updated: Aug 8, 2025

A Two-interval Forced-choice Task for Multisensory Comparisons
07:13

A Two-interval Forced-choice Task for Multisensory Comparisons

Published on: November 9, 2018

11.0K

Approaches to analyzing binary data for large-scale A/B testing.

Wenru Zhou1, Miranda Kroehl2, Maxene Meier3

  • 1Department of Biostatistics & Informatics, University of Colorado, United States.

Contemporary Clinical Trials Communications
|March 6, 2023
PubMed
Summary

The t-test performs reliably for binary outcomes in large-scale A/B testing, even with interim analyses. Naïve interim monitoring without corrections significantly degrades study performance.

Keywords:
A/B testingAcademic-industry partnershipInterim monitoringO'Brien-Fleming boundaries

More Related Videos

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K
Methods for Presenting Real-world Objects Under Controlled Laboratory Conditions
06:54

Methods for Presenting Real-world Objects Under Controlled Laboratory Conditions

Published on: June 21, 2019

6.0K

Related Experiment Videos

Last Updated: Aug 8, 2025

A Two-interval Forced-choice Task for Multisensory Comparisons
07:13

A Two-interval Forced-choice Task for Multisensory Comparisons

Published on: November 9, 2018

11.0K
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K
Methods for Presenting Real-world Objects Under Controlled Laboratory Conditions
06:54

Methods for Presenting Real-world Objects Under Controlled Laboratory Conditions

Published on: June 21, 2019

6.0K

Area of Science:

  • Statistics
  • Experimental Design
  • Industry Applications

Background:

  • Industry standard uses t-tests for all A/B testing outcomes (continuous and binary).
  • Naïve interim monitoring strategies are often applied without evaluating impact on power and type I error rates.
  • Robustness of t-tests is known, but performance with large-scale binary A/B testing data and interim analyses requires specific evaluation.

Purpose of the Study:

  • Evaluate statistical tests and study designs for large-scale industry A/B testing.
  • Assess the t-test's performance on binary outcomes with and without interim analyses.
  • Compare naïve interim monitoring with corrected approaches (e.g., O'Brien-Fleming).

Main Methods:

  • Simulation studies comparing t-test, Chi-squared test, and Chi-squared test with Yate's correction for binary data.
  • Evaluation of interim monitoring strategies: naïve approach vs. O'Brien-Fleming boundary.
  • Consideration of early termination for futility, difference, or both.

Main Results:

  • The t-test demonstrates comparable power and type I error rates for binary outcomes in large-scale industrial A/B tests.
  • Performance is consistent with and without interim monitoring.
  • Naïve interim monitoring without multiple testing corrections results in poorly performing studies.

Conclusions:

  • The t-test is a robust choice for binary outcomes in large-scale industrial A/B testing.
  • Interim analyses do not negatively impact the t-test's performance with large sample sizes.
  • Corrected interim monitoring is crucial; naïve approaches should be avoided to maintain study integrity.