Related Experiment Video
Updated: Aug 5, 2026

Optimization for Sequencing and Analysis of Degraded FFPE-RNA Samples
Published on: June 8, 2020
Modular RNA-Sequencing Analytics for Exploratory Biomarker Discovery Using Public Data
Cheryl L Sesler1, Lukasz S Wylezinski2, Guzel I Shaginurova1
1Decode Health, Inc., Nashville, Tennessee.
Abstract:
Publicly available RNA-sequencing data provide a cost-effective resource for biomarker discovery. However, heterogeneity across studies often complicates analysis. This study presents a modular analytics pipeline that combines publicly available data sets with established open-source tools; standardizes quality control, differential expression analysis, and pathway analysis; and leverages competitive machine learning to unify disparate RNA-sequencing data sets for robust biomarker identification. The workflow is demonstrated across three disease contexts, ranging from small pilot data sets to larger, integrated analyses: i) identifying differentially expressed gene signatures associated with coronavirus disease 2019 (COVID-19) severity, ii) combining differential expression and machine learning techniques to analyze multicohort sepsis data sets, resulting in concise biomarker panels, and iii) using both bulk and single-cell data to examine tissue and cell type specificity of N-acyl-phosphatidylethanolamine phospholipase D in atherosclerosis. These applications demonstrate how an adaptable, modular pipeline using open-source tools can repurpose public data to reduce noise, generate new hypotheses, and reveal meaningful biological insights, thereby establishing a foundation for future research and highlighting the importance of public data for exploratory biomarker discovery.
