Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Extraction: Partition and Distribution Coefficients01:14

Extraction: Partition and Distribution Coefficients

The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an organic...
Aggregates Classification01:29

Aggregates Classification

Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
How Data are Classified: Numerical Data00:59

How Data are Classified: Numerical Data

Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
How Data are Classified: Categorical Data01:11

How Data are Classified: Categorical Data

A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Propensity Score-Based Stratified Win Ratio for Augmented Control Designs.

Statistics in medicine·2026
Same author

Bayesian doubly robust estimation of causal effects for clustered observational data.

Journal of applied statistics·2025
Same author

Examining the protective effects of caregiver-child closeness on the association between parenting behaviors and youth aggression.

Scientific reports·2025
Same author

Aditus ad antrum patency on CT as a predictor of tympanoplasty outcomes in chronic otitis media.

Scientific reports·2025
Same author

Frequency-to-Place Mismatch and Cochlear Implant Outcomes-Beyond Electrode Type.

JAMA otolaryngology-- head & neck surgery·2025
Same author

High glucose levels promote glycolysis and cholesterol synthesis via ERRα and suppress the autophagy-lysosomal pathway in endometrial cancer.

Cell death & disease·2025

Related Experiment Video

Updated: Jun 20, 2026

A Method for Investigating Age-related Differences in the Functional Connectivity of Cognitive Control Networks Associated with Dimensional Change Card Sort Performance
09:01

A Method for Investigating Age-related Differences in the Functional Connectivity of Cognitive Control Networks Associated with Dimensional Change Card Sort Performance

Published on: May 7, 2014

Classification for high-throughput data with an optimal subset of principal components.

Joon Jin Song1, Yuan Ren, Fenglan Yan

  • 1Department of Mathematical Sciences, University of Arkansas, Fayetteville, AR 72701, USA. jjsong@uark.edu

Computational Biology and Chemistry
|September 15, 2009
PubMed
Summary

This study introduces a new ranking criterion for principal components (PCs) in high-throughput data analysis. This canonical variate criterion improves classification accuracy compared to traditional eigenvalue-based ranking.

More Related Videos

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

An Introduction to Processing, Fitting, and Interpreting Transient Absorption Data
08:12

An Introduction to Processing, Fitting, and Interpreting Transient Absorption Data

Published on: February 16, 2024

Related Experiment Videos

Last Updated: Jun 20, 2026

A Method for Investigating Age-related Differences in the Functional Connectivity of Cognitive Control Networks Associated with Dimensional Change Card Sort Performance
09:01

A Method for Investigating Age-related Differences in the Functional Connectivity of Cognitive Control Networks Associated with Dimensional Change Card Sort Performance

Published on: May 7, 2014

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

An Introduction to Processing, Fitting, and Interpreting Transient Absorption Data
08:12

An Introduction to Processing, Fitting, and Interpreting Transient Absorption Data

Published on: February 16, 2024

Area of Science:

  • Bioinformatics and Computational Biology
  • Statistical Genomics
  • Cheminformatics

Background:

  • High-throughput data are crucial for discovering gene and protein functions in biological and medical research.
  • Principal Component Analysis (PCA) is commonly used for dimensionality reduction in high-throughput data.
  • Traditional PCA component ranking by eigenvalues (variance) may not optimize subsequent multivariate analyses, especially classification.

Purpose of the Study:

  • To propose and evaluate an alternative ranking criterion for principal components (PCs) to enhance classification performance.
  • To compare the effectiveness of the canonical variate criterion against the classical eigenvalue criterion for dimensionality reduction.
  • To assess the performance of four prevalent classification methods using the proposed criterion on real-world high-throughput datasets.

Main Methods:

  • Application of the canonical variate criterion, focusing on within- and between-group variance, for ranking principal components.
  • Comparison with the classical eigenvalue-based ranking criterion.
  • Utilized four classification methods and leave-one-out cross-validation for performance evaluation.
  • Tested on three diverse high-throughput datasets: two microarray and one nuclear magnetic resonance (NMR) spectra dataset.

Main Results:

  • The canonical variate criterion demonstrated improved performance in classification tasks compared to the traditional eigenvalue-based ranking.
  • The proposed method effectively maximizes relevant information from a subset of principal components for better discrimination.
  • Consistent improvements were observed across different classification algorithms and diverse high-throughput data types.

Conclusions:

  • The canonical variate criterion offers a superior approach for selecting principal components in high-throughput data analysis, particularly for classification.
  • This method enhances the utility of PCA by prioritizing components that best differentiate between groups.
  • The findings suggest a more effective strategy for dimensionality reduction in complex biological and spectral data analysis.