Jove
Visualize
Contact Us

Related Concept Videos

Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

3.4K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.4K
Stratified Sampling Method01:16

Stratified Sampling Method

14.4K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
14.4K
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

6.8K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.8K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data01:16

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data

411
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
411
Test for Homogeneity01:23

Test for Homogeneity

2.3K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.3K
Surveys02:16

Surveys

16.6K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
16.6K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Estimating the Effect of Hospital Admission on Health Care Outcomes and Spending Among Persons With Dementia : A Quasi-experimental Study.

Annals of internal medicine·2026
Same author

Estimating the effects of hypothetical loneliness interventions on memory function among middle-aged and older adults in the United States.

American journal of epidemiology·2026
Same author

Explaining women's skepticism toward artificial intelligence: The role of risk orientation and risk exposure.

PNAS nexus·2026
Same author

Why publishing referee reports could backfire on public trust.

Nature·2025
Same author

How do the effects of toxicity in competitive online video games vary by source and match outcome?

PloS one·2025
Same author

Partisanship overcomes framing in shaping solar geoengineering perceptions: Evidence from a conjoint experiment.

npj climate action·2025
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jan 4, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

7.3K

Using Machine Learning to Uncover Hidden Heterogeneities in Survey Data.

Christina M Ramirez1, Marisa A Abrajano2, R Michael Alvarez3

  • 1Department of Biostatistics, UCLA Fielding School of Public Health, UCLA, Los Angeles, CA, 90095-1772, USA. cr@g.ucla.edu.

Scientific Reports
|November 7, 2019
PubMed
Summary

Public health survey language impacts response quality. Non-English responses show significant differences in health outcomes, highlighting the need for careful analysis in diverse populations.

More Related Videos

Constructing and Visualizing Models using Mime-based Machine-learning Framework
06:19

Constructing and Visualizing Models using Mime-based Machine-learning Framework

Published on: July 22, 2025

2.1K
Author Spotlight: Automated Lifespan Monitoring – Discovering Aging Dynamics with the Lifespan Machine
08:53

Author Spotlight: Automated Lifespan Monitoring – Discovering Aging Dynamics with the Lifespan Machine

Published on: January 26, 2024

1.6K

Related Experiment Videos

Last Updated: Jan 4, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

7.3K
Constructing and Visualizing Models using Mime-based Machine-learning Framework
06:19

Constructing and Visualizing Models using Mime-based Machine-learning Framework

Published on: July 22, 2025

2.1K
Author Spotlight: Automated Lifespan Monitoring – Discovering Aging Dynamics with the Lifespan Machine
08:53

Author Spotlight: Automated Lifespan Monitoring – Discovering Aging Dynamics with the Lifespan Machine

Published on: January 26, 2024

1.6K

Area of Science:

  • Public Health
  • Health Services Research
  • Computational Linguistics

Background:

  • Survey responses in public health are known to be heterogeneous.
  • Factors influencing response quality include cognitive abilities, interview context, and mode of administration.
  • The impact of language on survey response quality remains underexplored.

Purpose of the Study:

  • To investigate the association between language used in public health surveys and survey response quality.
  • To introduce and apply a machine learning approach, Fuzzy Forests, for analyzing survey response heterogeneity.
  • To assess differences in reported health outcomes based on language, including non-English and Asian languages.

Main Methods:

  • Utilized the 2013 California Health Interview Survey (CHIS) as a training dataset.
  • Employed the 2014 CHIS as a test dataset to validate the model.
  • Applied the Fuzzy Forests machine learning methodology for model selection and prediction.

Main Results:

  • Non-English language survey responses differed substantially from English responses in reported health outcomes.
  • Significant heterogeneity was observed among Asian languages, necessitating caution in cross-linguistic comparisons.
  • The Fuzzy Forests model accurately predicted 86% of good health outcomes using the 2014 test data.

Conclusions:

  • Language is a critical factor influencing survey response heterogeneity in public health.
  • The Fuzzy Forests methodology offers a promising tool for identifying and understanding diverse sources of survey response variation.
  • Findings underscore the importance of considering linguistic factors in the design and interpretation of public health surveys, especially complex ones.