Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Types of Selection01:46

Types of Selection

46.2K
Natural selection influences the frequencies of particular alleles and phenotypes within populations in several different ways. Primarily, natural selection can be directional, stabilizing, or disruptive. Directional selection favors one extreme trait and shifts the population towards that phenotype while selecting against individuals displaying alternate traits. Stabilizing selection favors an intermediate trait with a narrow range of variation. Deviation from the optimal phenotype towards an...
46.2K
Aggregates Classification01:29

Aggregates Classification

1.1K
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
1.1K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data01:16

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data

570
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
570
Survival Tree01:19

Survival Tree

464
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
464
Frequency-dependent Selection01:21

Frequency-dependent Selection

24.4K
When the fitness of a trait is influenced by how common it is (i.e., its frequency) relative to different traits within a population, this is referred to as frequency-dependent selection. Frequency-dependent selection may occur between species or within a single species. This type of selection can either be positive—with more common phenotypes having higher fitness—or negative, with rarer phenotypes conferring increased fitness.
24.4K
Classification of Systems-II01:31

Classification of Systems-II

547
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
547

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

High protective efficacy of a recombinant Fer2 vaccine against Dermacentor marginatus infestations.

Experimental & applied acarology·2026
Same author

Anomalous Saturation of CO Adsorption at 26% on Cu(111) Governed by Nanometer-Scale Substrate-Mediated Interactions.

Journal of the American Chemical Society·2025
Same author

Opportunities and challenges of diffusion models for generative AI.

National science review·2024
Same author

Are Latent Factor Regression and Sparse Regression Adequate?

Journal of the American Statistical Association·2024
Same author

Understanding Implicit Regularization in Over-Parameterized Single Index Model.

Journal of the American Statistical Association·2024
Same author

Communication-Efficient Accurate Statistical Estimation.

Journal of the American Statistical Association·2023

Related Experiment Video

Updated: Mar 21, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

8.1K

Feature Augmentation via Nonparametrics and Selection (FANS) in High-Dimensional Classification.

Jianqing Fan1, Yang Feng2, Jiancheng Jiang3

  • 1Jianqing Fan is Frederick L. Moore Professor of Finance, Department of Operations Research and Financial Engineering, Princeton University, Princeton, NJ, 08544 ( jqfan@princeton.edu ).

Journal of the American Statistical Association
|May 18, 2016
PubMed
Summary

We introduce Feature Augmentation via Nonparametrics and Selection (FANS), a high-dimensional classification method. FANS uses nonparametric feature augmentation and penalized logistic regression for competitive performance on complex data.

Keywords:
classificationdensity estimationfeature augmentationfeature selectionhigh dimensional spacenonlinear decision boundaryparallel computing

More Related Videos

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.4K

Related Experiment Videos

Last Updated: Mar 21, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

8.1K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.4K

Area of Science:

  • Statistics
  • Machine Learning
  • Computational Biology

Background:

  • High-dimensional data presents challenges for traditional classification methods.
  • Marginal density ratios are powerful univariate classifiers.
  • Existing methods may struggle with the curse of dimensionality and lack interpretability.

Purpose of the Study:

  • To propose a novel high-dimensional classification method called Feature Augmentation via Nonparametrics and Selection (FANS).
  • To develop a flexible nonlinear decision boundary that avoids the curse of dimensionality.
  • To enhance model interpretability and computability compared to related models.

Main Methods:

  • Nonparametric feature augmentation using marginal density ratio estimates.
  • Transformation of original feature measurements.
  • Application of penalized logistic regression on augmented features.
  • Generalization of the Naive Bayes model.

Main Results:

  • FANS creates models with local complexity and global simplicity.
  • Risk bounds for FANS are theoretically developed.
  • Numerical analysis shows FANS is competitive with existing methods.
  • Real-world data analysis on email spam and gene expression datasets demonstrates strong performance.

Conclusions:

  • FANS offers a powerful and flexible approach to high-dimensional classification.
  • The method is computationally efficient, utilizing parallel computing.
  • FANS provides a competitive alternative for analyzing complex datasets in various domains.