Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Prediction Intervals01:03

Prediction Intervals

2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
2.3K
Assumptions of Survival Analysis01:15

Assumptions of Survival Analysis

126
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
126
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data01:16

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data

129
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
129
Relative Risk01:12

Relative Risk

163
Relative risk (RR) is a statistical measure commonly used in epidemiology to compare the likelihood of a particular event occurring between two groups. This metric is important for evaluating the relationship between exposure to a specific risk factor and the probability of a particular outcome. It plays a crucial role in medical research, public health studies, and risk assessment. Relative risk quantifies how much more (or less) likely an event is to occur in an exposed group compared to an...
163
Multiple Regression01:25

Multiple Regression

3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches01:23

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches

126
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
126

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A rare case of intramuscular granular cell tumor in the right thigh: case report and literature review.

Frontiers in medicine·2026
Same author

Cancer risk with methotrexate-TNF antagonist combination therapy compared to TNF-antagonist monotherapy in IBD.

Inflammatory bowel diseases·2026
Same author

Toward Artificial Intelligence-driven Clinical Decision Support Tools in Rheumatology.

Rheumatic diseases clinics of North America·2026
Same author

A novel <i>ABO</i> splice site variant underlying the A<sub>3</sub> phenotype: immunogenetic basis and functional dissection.

Frontiers in genetics·2026
Same author

A magnetic resonance imaging-guided drug delivery system for premetastatic niche theranostics and colorectal cancer liver metastasis intervention.

Journal of nanobiotechnology·2026
Same author

Statistics and AI - A Fireside Conversation.

Harvard data science review·2026

Related Experiment Video

Updated: Jun 30, 2025

Establishing a Competing Risk Regression Nomogram Model for Survival Data
04:57

Establishing a Competing Risk Regression Nomogram Model for Survival Data

Published on: October 23, 2020

10.2K

Surrogate Assisted Semi-supervised Inference for High Dimensional Risk Prediction.

Jue Hou1, Zijian Guo2, Tianxi Cai3

  • 1Division of Biostatistics, University of Minnesota School of Public Health, Minneapolis, MN 55455, USA.

Journal of Machine Learning Research : JMLR
|March 19, 2024
PubMed
Summary

This study introduces a novel semi-supervised learning method for risk modeling using electronic health records (EHR). The approach effectively uses unlabeled data to improve disease risk prediction, even with missing outcome information.

Keywords:
generalized linear modelshigh dimensional inferencemodel mis-specificationrisk predictionsemi-supervised learning

More Related Videos

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.1K
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

15.7K

Related Experiment Videos

Last Updated: Jun 30, 2025

Establishing a Competing Risk Regression Nomogram Model for Survival Data
04:57

Establishing a Competing Risk Regression Nomogram Model for Survival Data

Published on: October 23, 2020

10.2K
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.1K
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

15.7K

Area of Science:

  • Biostatistics
  • Machine Learning
  • Health Informatics

Background:

  • Risk modeling using electronic health records (EHR) is hindered by unobserved disease outcomes and high-dimensional predictors.
  • Existing methods struggle with incomplete outcome data and complex predictor sets.

Purpose of the Study:

  • To develop a robust semi-supervised learning approach for risk prediction using EHR data.
  • To address challenges of missing outcomes and high-dimensional predictors in EHR-based risk modeling.
  • To leverage both labeled and unlabeled data for improved statistical inference.

Main Methods:

  • A surrogate-assisted semi-supervised learning framework is proposed.
  • Unobserved outcomes are imputed using a sparse imputation model with outcome surrogates and high-dimensional predictors.
  • A one-step bias correction is applied for valid interval estimation in risk prediction.

Main Results:

  • The proposed method demonstrates superiority over existing supervised methods in extensive simulation studies.
  • The inference procedure remains valid even with misspecified imputation and risk prediction models.
  • The approach enables high-dimensional statistical inference for dense risk prediction models.

Conclusions:

  • The novel approach effectively utilizes unlabeled EHR data for enhanced disease risk prediction.
  • This method offers a powerful tool for genetic risk prediction, as demonstrated in a type-2 diabetes mellitus cohort.
  • The technique provides a valid and robust framework for statistical inference in challenging EHR data settings.