Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Variation01:19

Variation

7.5K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
7.5K
Mean Absolute Deviation01:13

Mean Absolute Deviation

3.1K
The mean absolute deviation is also a measure of the variability of data in a sample. It is the absolute value of the average difference between the data values and the mean.
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
3.1K
Weighted Mean00:57

Weighted Mean

6.0K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
6.0K
Variation: Normal Distribution, Range, and Standard Deviation02:32

Variation: Normal Distribution, Range, and Standard Deviation

25.3K
In the field of psychology, there are several ways to organize measurements of a trait, feature, or characteristic (i.e., variables). Qualitative data, such as ethnicity, can be tabulated into a frequency count to provide information about the proportion, as well as the variety of groups in a sample or population. On the other hand, researchers can perform a wider set of calculations on quantitative data. The mean, mode, and median, for instance, are central tendency measures to identify a...
25.3K
Testing a Claim about Standard Deviation01:19

Testing a Claim about Standard Deviation

2.7K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.7K
Chebyshev's Theorem to Interpret Standard Deviation01:15

Chebyshev's Theorem to Interpret Standard Deviation

4.8K
Chebyshev’s theorem, also known as Chebyshev’s Inequality, states that the proportion of values of a dataset for K standard deviation is calculated using the equation:
4.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Factors Influencing Breastfeeding Knowledge and Attitudes Among Parents of Moderate-To-Late Preterm Infants: A Cross-Sectional Study.

Advances in neonatal care : official journal of the National Association of Neonatal Nurses·2026
Same author

Chitosan-based thermosensitive hydrogel for nasal delivery of exenatide: Effect of magnesium chloride.

International journal of pharmaceutics·2018
Same author

Surface-functionalized, pH-responsive poly(lactic-co-glycolic acid)-based microparticles for intranasal vaccine delivery: Effect of surface modification with chitosan and mannan.

European journal of pharmaceutics and biopharmaceutics : official journal of Arbeitsgemeinschaft fur Pharmazeutische Verfahrenstechnik e.V·2016
Same author

Stabilization and immune response of HBsAg encapsulated within poly(lactic-co-glycolic acid) microspheres using HSA as a stabilizer.

International journal of pharmaceutics·2015
Same author

Formulation and evaluation of poly(lactic-co-glycolic acid) microspheres loaded with an altered collagen type II peptide for the treatment of rheumatoid arthritis.

Journal of microencapsulation·2015
Same author

Effect of site-specific PEGylation on the fibrinolytic activity, immunogenicity, and pharmacokinetics of staphylokinase.

Acta biochimica et biophysica Sinica·2014

Related Experiment Video

Updated: Nov 27, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.8K

Weighted Mean Squared Deviation Feature Screening for Binary Features.

Gaizhen Wang1, Guoyu Guan2

  • 1School of Mathematics and Statistics, Northeast Normal University, Changchun 130000, China.

Entropy (Basel, Switzerland)
|December 8, 2020
PubMed
Summary

We introduce weighted mean squared deviation (WMSD), a new model-free method for selecting important binary features in ultrahigh dimensions. WMSD effectively identifies relevant features, especially those with probabilities near 0.5, outperforming existing methods.

Keywords:
Chi-square statisticPearson correlation coefficientfeature screeningmutual informationpower-law distributionweighted mean squared deviation

More Related Videos

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.1K
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

16.0K

Related Experiment Videos

Last Updated: Nov 27, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.8K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.1K
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

16.0K

Area of Science:

  • Machine Learning
  • Statistical Modeling
  • Data Mining

Background:

  • Ultrahigh dimensional data presents challenges for feature selection in binary classification.
  • Existing methods like Chi-square and mutual information may not optimally handle binary features with probabilities near 0.5.

Purpose of the Study:

  • To propose a novel model-free feature screening method for ultrahigh dimensional binary classification.
  • To address limitations of existing methods in identifying informative binary features.

Main Methods:

  • Introduced weighted mean squared deviation (WMSD) for feature screening.
  • Theoretically investigated asymptotic properties under the assumption log p = o(n).
  • Employed a Pearson correlation coefficient method for practical feature selection based on power-law distribution.

Main Results:

  • WMSD demonstrates superior performance compared to Chi-square and mutual information, particularly for features with probabilities near 0.5.
  • The method's theoretical properties were rigorously examined.
  • Empirical results on Chinese text classification show effectiveness with a small selected feature dimension.

Conclusions:

  • The proposed WMSD method is a powerful tool for feature screening in ultrahigh dimensional binary classification.
  • WMSD offers advantages in identifying relevant binary features, enhancing classification model performance.
  • The method is validated through theoretical analysis and practical application in text classification.