Related Experiment Video
Updated: Sep 2, 2025

A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types
Published on: December 10, 2012
Iterative bicluster-based Bayesian principal component analysis and least squares for missing-value imputation in
Saskya Mary Soemartojo1, Titin Siswantining1, Yoel Fernando1
1Department of Mathematics, Faculty of Mathematics and Natural Sciences, Universitas Indonesia, Indonesia.
A new gene expression imputation method, iterative bicluster-based Bayesian principal component analysis and least squares (bi-BPCA-iLS), improves accuracy over existing techniques. This method enhances missing value estimation in microarray and RNA-seq data with minimal extra computation.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Gene expression data from microarray and RNA-sequencing (RNA-seq) often contain missing values, necessitating imputation.
- Current methods like iterative bicluster-based least squares (bi-iLS) use biclustering but rely on row averages, which can misrepresent data structure.
- Row averages fail to capture the complex relationships within gene expression datasets.
Purpose of the Study:
- To introduce a novel missing-value imputation method, iterative bicluster-based Bayesian principal component analysis and least squares (bi-BPCA-iLS).
- To improve the accuracy of gene expression imputation by replacing row averages with Bayesian principal component analysis (BPCA) for temporary matrix completion.
- To evaluate the performance of bi-BPCA-iLS on both microarray and RNA-seq datasets.
Main Methods:
- Developed the iterative bicluster-based Bayesian principal component analysis and least squares (bi-BPCA-iLS) method.
- Utilized Bayesian principal component analysis (BPCA) to generate a temporary complete matrix, improving upon the row average method.
- Validated the method using yeast Saccharomyces cerevisiae microarray and Schizosaccharomyces pombe RNA-seq gene expression datasets.
Main Results:
- The bi-BPCA-iLS method demonstrated significant improvements in normalized root mean square error (NRMSE) for microarray data (10.6% vs. LLS, 0.6% vs. bi-iLS).
- For RNA-seq data, bi-BPCA-iLS showed notable NRMSE improvements (8.2% vs. LLS, 3.1% vs. bi-iLS).
- The computational cost of bi-BPCA-iLS was found to be comparable to the bi-iLS method.
Conclusions:
- The proposed bi-BPCA-iLS method offers a more accurate approach for imputing missing values in gene expression data.
- Replacing row averages with BPCA effectively captures dataset structure, leading to enhanced imputation performance.
- bi-BPCA-iLS provides a computationally efficient and accurate solution for handling missing data in both microarray and RNA-seq analyses.
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
RACE - Rapid Amplification of cDNA Ends

