Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

One-Way ANOVA: Equal Sample Sizes01:15

One-Way ANOVA: Equal Sample Sizes

4.0K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
4.0K
Chi-square Distribution01:10

Chi-square Distribution

6.5K
How does one determine if bingo numbers are evenly distributed or if some numbers occurred with a greater frequency? Or if the types of movies people preferred were different across different age groups or if a coffee machine dispensed approximately the same amount of coffee each time. These questions can be addressed by conducting a hypothesis test. One distribution that can be used to find answers to such questions is known as the chi-square distribution. The chi-square distribution has...
6.5K
One-Way ANOVA: Unequal Sample Sizes01:15

One-Way ANOVA: Unequal Sample Sizes

6.6K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
6.6K
Distributions to Estimate Population Parameter01:26

Distributions to Estimate Population Parameter

5.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
5.0K
Extraction: Partition and Distribution Coefficients01:14

Extraction: Partition and Distribution Coefficients

4.6K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
4.6K
Wilcoxon Signed-Ranks Test for Median of Single Population01:14

Wilcoxon Signed-Ranks Test for Median of Single Population

432
The Wilcoxon signed-rank test for the median of a single population is a nonparametric test used to evaluate whether the median of a population differs from a specified value. Unlike parametric tests, it does not require data to follow a normal distribution, making it suitable for non-normal or small samples. The test begins by calculating the difference (d) between each observation and the hypothesized median. The absolute values of these differences are ranked in ascending order, with ties...
432

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Random-with-constraints: Constructing minimal models for high-dimensional biology.

Proceedings of the National Academy of Sciences of the United States of America·2026
Same author

The hierarchical timescale hypothesis: Functional and structural convergence of biological networks and artificial neural nets.

Cell systems·2026
Same author

Randomness with constraints: constructing minimal models for high-dimensional biology.

ArXiv·2025
Same author

Physics-tailored machine learning reveals unexpected physics in dusty plasmas.

Proceedings of the National Academy of Sciences of the United States of America·2025
Same author

A mathematical model for ketosis-prone diabetes suggests the existence of multiple pancreatic β-cell inactivation mechanisms.

eLife·2025
Same author

Variance in C. elegans gut bacterial load suggests complex host-microbe dynamics.

PLoS computational biology·2025

Related Experiment Video

Updated: Jan 14, 2026

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

16.1K

Distribution of singular values in large sample cross-covariance matrices.

Arabind Swain1, Sean Alexander Ridout2, Ilya Nemenman3

  • 1Emory University, Department of Physics, Atlanta, Georgia 30322, USA.

Physical Review. E
|October 21, 2025
PubMed
Summary

This study introduces a new method to analyze cross-correlations in high-dimensional datasets, extending the Marchenko-Pastur theorem. The findings enable signal detection even when data dimensions exceed sample size, crucial for data science applications.

More Related Videos

Cross-Modal Multivariate Pattern Analysis
13:51

Cross-Modal Multivariate Pattern Analysis

Published on: November 9, 2011

20.4K
Basics of Multivariate Analysis in Neuroimaging Data
06:35

Basics of Multivariate Analysis in Neuroimaging Data

Published on: July 24, 2010

17.3K

Related Experiment Videos

Last Updated: Jan 14, 2026

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

16.1K
Cross-Modal Multivariate Pattern Analysis
13:51

Cross-Modal Multivariate Pattern Analysis

Published on: November 9, 2011

20.4K
Basics of Multivariate Analysis in Neuroimaging Data
06:35

Basics of Multivariate Analysis in Neuroimaging Data

Published on: July 24, 2010

17.3K

Area of Science:

  • Statistics
  • Data Science
  • Machine Learning

Background:

  • Estimating cross-covariance for high-dimensional datasets (N > T) is challenging due to large sampling fluctuations.
  • Existing methods like whitening fail when data dimensionality exceeds the number of samples.

Purpose of the Study:

  • To derive the probability distribution of singular values for empirical cross-covariances of high-dimensional datasets.
  • To extend the Marchenko-Pastur theorem to analyze cross-covariance matrices.
  • To enable signal detection in scenarios where traditional methods are inadequate.

Main Methods:

  • Analyzing uncorrelated Gaussian i.i.d. matrices X and Y with dimensions T×N_X and T×N_Y.
  • Deriving the probability distribution of singular values of XᵀY in various parameter regimes.
  • Investigating limiting cases of the derived distributions.

Main Results:

  • The probability distribution of singular values for empirical cross-covariances is derived.
  • This extends the Marchenko-Pastur result for sample covariance matrices.
  • The derived distributions allow for signal detection even when N > T.

Conclusions:

  • The new method provides a robust framework for analyzing cross-correlations in high-dimensional data.
  • It offers a way to establish statistical significance for cross-correlations, even in challenging N > T scenarios.
  • This research has broad implications for various data science applications requiring robust correlation analysis.