Related Experiment Video
Updated: Jun 6, 2026

Optimization for Sequencing and Analysis of Degraded FFPE-RNA Samples
Published on: June 8, 2020
Evaluating Cross-Platform Batch Correction Methods for Integrated Microarray and RNA-seq Data Analysis
Xuejun Sun1, Yu Zhang1, Chuwen Liu1
1Department of Biostatistics, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina, U.S.A.
Abstract:
Integrated analysis of gene expression data across studies is important for understanding complex traits, but combining microarray and RNA-seq data remains challenging because of substantial platform-related technical differences. In this study, we evaluated ten commonly used batch-effect correction methods for cross-platform integration and classified them into three groups representing a spectrum from least to most data processing: unsupervised sample-wise methods, unsupervised gene-wise methods, and supervised methods. Unlike the first two groups, supervised methods use outcome information to guide correction. Performance was assessed for distribution alignment, clustering, outcome prediction, and differential expression (DE) analysis using three paired real microarray-RNA-seq datasets and simulation studies. Supervised methods achieved strong distribution alignment and clustering by biological group, but they introduced information leakage and inflated Type I error, making them unsuitable for unbiased discovery. Among unsupervised methods, sample-wise approaches were often conservative under balanced settings. They showed inflated Type I error under imbalanced settings. In contrast, gene-wise methods generally maintained appropriate Type I error control and achieved higher power, with limma showing the strongest DE performance. Meta-analysis that combines p-values also controlled Type I error well and provided competitive power. For outcome prediction, both sample-wise QN and gene-wise limma performed well, with limma providing the strongest overall performance and QN offering the practical advantage that new samples can be normalized to an existing reference distribution without refitting. Overall, limma is recommended as the best general-purpose method for cross-platform integration.
