Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Steps in Outbreak Investigation01:18

Steps in Outbreak Investigation

411
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
411
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

792
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
792
Estimating Population Mean with Unknown Standard Deviation01:22

Estimating Population Mean with Unknown Standard Deviation

8.7K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
8.7K
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

8.7K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.7K
Estimating Population Standard Deviation01:26

Estimating Population Standard Deviation

3.2K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.2K
Estimating Population Mean with Known Standard Deviation01:16

Estimating Population Mean with Known Standard Deviation

9.5K
To construct a confidence interval for a single unknown population mean μ, where the population standard deviation is known, we need sample mean as an estimate for μ and we need the margin of error. Here, the margin of error (EBM) is called the error bound for a population mean (abbreviated EBM). The sample mean is the point estimate of the unknown population mean μ.
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
9.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

With Gratitude to Dr. Thomas Einhorn, Founding Editor of JBJS Reviews: Honoring a Legacy of Orthopaedic and Editorial Leadership.

The Journal of bone and joint surgery. American volume·2026
Same author

Association Between Complications and Death Within 30 Days After Orthopedic Surgery: Vascular Events in Noncardiac Surgery Patients Cohort Evaluation (VISION) Substudy.

JMIR perioperative medicine·2026
Same author

With Gratitude to Dr. Thomas Einhorn, Founding Editor of JBJS Reviews: Honoring a Legacy of Orthopaedic and Editorial Leadership.

The Journal of bone and joint surgery. American volume·2026
Same author

Musculoskeletal Manifestations of Perimenopause: A Systematic Review and Meta-Analysis of 93,021 Women.

JB & JS open access·2026
Same author

Beyond Words: Moving from Ideas to Action.

The Journal of bone and joint surgery. American volume·2025
Same author

Pain as a Driver of Myocardial Injury in Hip Fracture Patients: A Hip Fracture Accelerated Surgical Treatment and Care Track (HIP ATTACK) Trial Secondary Analysis.

Anesthesiology·2025

Related Experiment Video

Updated: Dec 16, 2025

A Data-Driven Approach to Quantifying Immune States in Sepsis
07:42

A Data-Driven Approach to Quantifying Immune States in Sepsis

Published on: February 7, 2025

414

Using Machine Learning to Estimate Unobserved COVID-19 Infections in North America.

Shashank Vaid1, Caglar Cakan2, Mohit Bhandari3

  • 1DeGroote School of Business, McMaster University, Hamilton, Ontario, Canada.

The Journal of Bone and Joint Surgery. American Volume
|July 4, 2020
PubMed
Summary

COVID-19 infections were significantly underestimated in North America as of April 2020. Machine learning models revealed millions of undetected cases in the United States and tens of thousands in Canada, highlighting challenges in disease detection.

More Related Videos

A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
10:46

A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data

Published on: December 9, 2015

10.9K
Large-Scale SARS-CoV-2 Testing Utilizing Saliva and Transposition Sample Pooling
08:26

Large-Scale SARS-CoV-2 Testing Utilizing Saliva and Transposition Sample Pooling

Published on: June 23, 2022

1.9K

Related Experiment Videos

Last Updated: Dec 16, 2025

A Data-Driven Approach to Quantifying Immune States in Sepsis
07:42

A Data-Driven Approach to Quantifying Immune States in Sepsis

Published on: February 7, 2025

414
A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
10:46

A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data

Published on: December 9, 2015

10.9K
Large-Scale SARS-CoV-2 Testing Utilizing Saliva and Transposition Sample Pooling
08:26

Large-Scale SARS-CoV-2 Testing Utilizing Saliva and Transposition Sample Pooling

Published on: June 23, 2022

1.9K

Area of Science:

  • Epidemiology
  • Machine Learning
  • Public Health

Background:

  • Coronavirus disease 2019 (COVID-19) detection remains a significant global challenge.
  • As of April 22, 2020, over 2.6 million infections and 183,000 deaths were reported worldwide.
  • Policymakers rely on projections for critical decisions, necessitating accurate case detection.

Purpose of the Study:

  • To model unobserved COVID-19 infections in North America.
  • To assess the extent of underestimation of COVID-19 cases.
  • To provide data-driven insights for public health policy.

Main Methods:

  • Developed a machine-learning model utilizing dimensionality reduction to identify key parameters.
  • Employed an unbiased hierarchical Bayesian estimator to infer past infections from current fatalities.
  • Analyzed reported cases and fatalities with varying lag times between infection and death.

Main Results:

  • The United States potentially had over 1.3 to 1.7 million undetected COVID-19 infections by April 22, 2020, depending on lag time assumptions.
  • Canada may have had 60,000 to 80,000 undetected infections during the same period.
  • These estimates significantly exceeded reported cases, indicating substantial undercounting.

Conclusions:

  • COVID-19 infections were likely 1.5 to 2.0 times higher than reported in the United States and 1.4 to 2.1 times higher in Canada.
  • Even with similar fatality and growth rates, undetected infections represent a substantial portion of the true burden.
  • Two distinct modeling approaches converged on similar estimates of undetected infections in North America.