Related Experiment Video
Updated: Sep 9, 2025

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.6K
Weighted Sparse Partial Least Squares With Joint Sample and Feature Selection for Integrating Multi-Omics Data
IEEE Transactions on Computational Biology and Bioinformatics
|September 3, 2025
Summary
This study introduces a new method for Sparse Partial Least Squares (sPLS) to identify specific sample subsets and remove outliers in data fusion. The novel approach enhances sPLS for improved multi-view data analysis and outlier detection.
Area of Science:
- Computational Biology
- Machine Learning
- Statistical Analysis
Background:
- Sparse Partial Least Squares (sPLS) is a dimensionality reduction technique for data fusion.
- Standard sPLS cannot identify latent subsets of samples or remove outliers.
Purpose of the Study:
- To develop a novel method for joint sample and feature selection in sPLS.
- To extend sPLS for identifying specific sample subsets and outlier removal.
- To adapt the method for multi-view data fusion.
Main Methods:
- Proposed an $\ell _\infty /\ell _{0}$-norm constrained weighted sparse PLS ($\ell _\infty /\ell _{0}$-wsPLS) for sample and feature selection.
- Proved the Kurdyka-Łojasiewicz property of the $\ell _\infty /\ell _{0}$-norm constraints for global convergence.
- Developed two multi-view wsPLS models and efficient iterative algorithms for multi-view data fusion.
Main Results:
- The proposed $\ell _\infty /\ell _{0}$-wsPLS method enables joint sample and feature selection.
- Globally convergent algorithms were developed for the proposed models.
- Numerical and biomedical data experiments demonstrated the efficiency of the multi-view wsPLS methods.
Conclusions:
- The novel $\ell _\infty /\ell _{0}$-wsPLS method effectively identifies sample subsets and outliers.
- The extended multi-view wsPLS models are efficient for multi-view data fusion.
- The developed algorithms ensure convergence and demonstrate practical applicability.
More Related Videos
Related Concept Videos
Cluster Sampling Method
12.7K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.7K
Stratified Sampling Method
12.8K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
To choose a stratified sample, divide the population into groups called strata and then take a...
12.8K
Multi-species Conserved Sequences
4.3K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.3K
Sampling Plans
258
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
258
Single Nucleotide Polymorphisms-SNPs
15.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.8K
Genome-wide Association Studies-GWAS
14.1K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.1K

