Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Causality in Epidemiology01:21

Causality in Epidemiology

863
Causality or causation is a fundamental concept in epidemiology, vital for understanding the relationships between various factors and health outcomes. Despite its importance, there's no single, universally accepted definition of causality within the discipline. Drawing from a systematic review, causality in epidemiology encompasses several definitions, including production, necessary and sufficient, sufficient-component, counterfactual, and probabilistic models. Each has its strengths and...
863
Variability: Analysis01:11

Variability: Analysis

191
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
191
Friedman Two-way Analysis of Variance by Ranks01:21

Friedman Two-way Analysis of Variance by Ranks

299
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
299
Correlation and Causation01:27

Correlation and Causation

39.6K
Statistical tests can calculate whether there is a relationship, or correlation, between independent and dependent variables. An indirect relationship of the variables signifies a correlation, while a direct relationship shows causation. If it is determined that no connection exists between the variables, then the correlation is a coincidence.
Correlation versus Causation
If the dependent variable increases or decreases when the independent variable increases, there is a positive or negative...
39.6K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

539
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
539
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data01:16

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data

215
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
215

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same authorSame Topic

Federated feature selection with false discovery rate control.

Journal of the Royal Statistical Society. Series B, Statistical methodology·2026
Same author

PDA (Privacy-Preserving Distributed Algorithms) in action: ten principles for high-quality multi-site clinical evidence generation.

Journal of the American Medical Informatics Association : JAMIA·2026
Same author

Promoting Problem-Solving Among Low-Income Adults With Type 2 Diabetes: Cluster-Randomized Controlled Trial of a Mobile Health Intervention With SMS Text Messaging (Mobile Diabetes Detective).

Journal of medical Internet research·2026
Same author

Real-world performance of large-scale propensity score adjustment strategies: Matching, weighting, and stratification.

Research square·2026
Same author

A comparison of Fast Healthcare Interoperability Resources and Observational Medical Mutcomes Partnership electronic health record data within the All of Us Research Program.

Journal of the American Medical Informatics Association : JAMIA·2026
Same author

Unlocking multi-institutional insights into disease progression with PEAL as a lossless, one-shot federated learning solution.

NPJ digital medicine·2026

Related Experiment Video

Updated: Sep 15, 2025

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

15.8K

DisC2o-HD: Distributed causal inference with covariates shift for analyzing real-world high-dimensional data.

Jiayi Tong1, Jie Hu1, George Hripcsak2

  • 1Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania, Philadelphia, PA 19104, USA.

Journal of Machine Learning Research : JMLR
|July 18, 2025
PubMed
Summary

This study introduces DisC²o-HD, a distributed learning algorithm for high-dimensional healthcare data. It effectively estimates average treatment effects (ATE) while addressing covariate shift across multiple clinical sites.

Keywords:
Causal InferenceDistribution ShiftFederated LearningHigh-dimensional DataReal-World Data

More Related Videos

Basics of Multivariate Analysis in Neuroimaging Data
06:35

Basics of Multivariate Analysis in Neuroimaging Data

Published on: July 24, 2010

17.0K
Application of Granger Causality Analysis of the Directed Functional Connection in Alzheimer's Disease and Mild Cognitive Impairment
08:43

Application of Granger Causality Analysis of the Directed Functional Connection in Alzheimer's Disease and Mild Cognitive Impairment

Published on: August 7, 2017

8.0K

Related Experiment Videos

Last Updated: Sep 15, 2025

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

15.8K
Basics of Multivariate Analysis in Neuroimaging Data
06:35

Basics of Multivariate Analysis in Neuroimaging Data

Published on: July 24, 2010

17.0K
Application of Granger Causality Analysis of the Directed Functional Connection in Alzheimer's Disease and Mild Cognitive Impairment
08:43

Application of Granger Causality Analysis of the Directed Functional Connection in Alzheimer's Disease and Mild Cognitive Impairment

Published on: August 7, 2017

8.0K

Area of Science:

  • Health Informatics
  • Biostatistics
  • Machine Learning

Background:

  • High-dimensional healthcare data (EHR, claims) present challenges: numerous variables, multi-site data consolidation, and covariate shift.
  • Estimating treatment effects in such data requires robust methods to handle heterogeneity across sites.

Purpose of the Study:

  • To propose a novel distributed learning algorithm, DisC²o-HD, for estimating the average treatment effect (ATE) in high-dimensional healthcare data.
  • To address covariate shift and data heterogeneity across multiple clinical sites.

Main Methods:

  • Developed DisC²o-HD, a distributed learning algorithm utilizing surrogate likelihood.
  • Employs propensity score and outcome model calibration to achieve covariate balancing and account for covariate shift.
  • Demonstrates that the distributed estimator approximates the pooled estimator.

Main Results:

  • The proposed estimator is consistent if either the propensity score or outcome regression model is correctly specified.
  • Achieves semiparametric efficiency when both models are correctly specified.
  • Simulation studies and real-world data application validate the algorithm's performance and readiness.

Conclusions:

  • DisC²o-HD offers a valid and implementable solution for estimating ATE in distributed, high-dimensional healthcare data with covariate shift.
  • The algorithm provides a robust approach to leveraging multi-site data while maintaining statistical validity.