Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Statistical Analysis: Overview01:11

Statistical Analysis: Overview

6.8K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.8K
Mechanistic Models: Compartment Models in Individual and Population Analysis01:23

Mechanistic Models: Compartment Models in Individual and Population Analysis

70
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
70
One-Way ANOVA01:18

One-Way ANOVA

8.1K
One-way ANOVA analyzes more than three samples categorized by one factor. For example, it can compare the average mileage of sports bikes. Here, the data is categorized by one factor - the company. However, one-way ANOVA cannot be used to simultaneously compare the sample mean of three or more samples categorized by two factors. An example of two factors would be sports bikes from different companies driven in different terrains, such as a desert or snowy landscape. Here, two-way ANOVA is used...
8.1K
Review and Preview01:10

Review and Preview

7.6K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
7.6K
Biostatistics: Overview01:20

Biostatistics: Overview

301
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
301
Statistical Methods to Analyze Parametric Data: ANOVA01:12

Statistical Methods to Analyze Parametric Data: ANOVA

497
Analysis of Variance, or ANOVA, is a powerful statistical technique used to analyze parametric data, primarily in research and experimental studies. It's designed to compare the means of two or more groups, assisting researchers in identifying any significant differences between these group means. There are two main types of ANOVA based on the complexity of the analysis: one-way and two-way.
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
497

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Metabolomic profiling of genotype-derived ABO blood group, secretor status, and Lewis antigens and association with pancreatic ductal adenocarcinoma risk.

American journal of epidemiology·2026
Same author

Mammographic Density in Relation to Breast Cancer Risk Factors among Chinese Women.

Cancer epidemiology, biomarkers & prevention : a publication of the American Association for Cancer Research, cosponsored by the American Society of Preventive Oncology·2024
Same author

A comprehensive framework for trans-ancestry pathway analysis using GWAS summary data from diverse populations.

PLoS genetics·2024
Same author

Bias and mean squared error in Mendelian randomization with invalid instrumental variables.

Genetic epidemiology·2023
Same author

Improve the model of disease subtype heterogeneity by leveraging external summary data.

PLoS computational biology·2023
Same author

Genetic Susceptibility to Nonalcoholic Fatty Liver Disease and Risk for Pancreatic Cancer: Mendelian Randomization.

Cancer epidemiology, biomarkers & prevention : a publication of the American Association for Cancer Research, cosponsored by the American Society of Preventive Oncology·2023

Related Experiment Video

Updated: Aug 5, 2025

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
08:51

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts

Published on: September 20, 2024

1.4K

Integrative analysis of individual-level data and high-dimensional summary statistics.

Sheng Fu1, Lu Deng2, Han Zhang3

  • 1Division of Cancer Epidemiology and Genetics, National Cancer Institute, Bethesda, MD 20892, USA.

Bioinformatics (Oxford, England)
|March 25, 2023
PubMed
Summary

This study introduces a novel procedure for efficient statistical inference by integrating individual-level data with high-dimensional summary statistics from genome-wide association studies. The method enhances effect estimation and hypothesis testing, offering a scalable solution for complex genetic data analysis.

More Related Videos

ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
05:12

ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data

Published on: January 16, 2019

11.5K
Profiling Maternal Behavior Responses During Whole-Brain Imaging
07:12

Profiling Maternal Behavior Responses During Whole-Brain Imaging

Published on: January 24, 2025

858

Related Experiment Videos

Last Updated: Aug 5, 2025

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
08:51

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts

Published on: September 20, 2024

1.4K
ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
05:12

ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data

Published on: January 16, 2019

11.5K
Profiling Maternal Behavior Responses During Whole-Brain Imaging
07:12

Profiling Maternal Behavior Responses During Whole-Brain Imaging

Published on: January 24, 2025

858

Area of Science:

  • Statistical Genetics
  • Bioinformatics
  • Computational Biology

Background:

  • Statistical analyses typically rely on individual-level data.
  • Integrating high-dimensional summary statistics (e.g., from genome-wide association studies) with individual data presents computational challenges.
  • Existing methods struggle with numeric issues when optimizing objective functions with many parameters.

Purpose of the Study:

  • To develop a procedure for efficient statistical inference by leveraging external summary data.
  • To enhance effect estimation and hypothesis testing using combined individual and summary-level data.
  • To create a scalable method for high-dimensional genetic data integration.

Main Methods:

  • A divide-and-conquer strategy is proposed to handle high-dimensional summary data by breaking down the task into parallel jobs.
  • Each job integrates individual-level data with a subset of summary data.
  • Final parameter estimates are obtained by pooling results from fitted models using minimum distance estimation.

Main Results:

  • The procedure improves the fitting of targeted statistical models, enhancing inference efficiency.
  • The method is scalable to high-dimensional summary data.
  • Demonstrated advantages through simulations and an application to pancreatic cancer risk prediction using polygenic risk scores.

Conclusions:

  • The developed procedure offers a robust and efficient approach for integrating individual-level and high-dimensional summary data.
  • The method is applicable to general additive models in genetic studies and can integrate data from different populations.
  • An R package (MetaGIM) is available for implementation.