Related Experiment Video
Updated: Aug 5, 2026

07:30
Optimization for Sequencing and Analysis of Degraded FFPE-RNA Samples
Published on: June 8, 2020
Modular RNA-Sequencing Analytics for Exploratory Biomarker Discovery Using Public Data
Cheryl L Sesler1, Lukasz S Wylezinski2, Guzel I Shaginurova1
1Decode Health, Inc., Nashville, Tennessee.
The Journal of Molecular Diagnostics : JMD
|July 31, 2026
Summary
This study introduces a modular pipeline to analyze diverse public RNA sequencing data for biomarker discovery. It integrates datasets and machine learning to overcome heterogeneity, enabling robust identification of disease-related genes.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Publicly available RNA sequencing (RNA-seq) data are valuable for biomarker discovery but suffer from cross-study heterogeneity.
- Analyzing disparate RNA-seq datasets requires robust methods to standardize data and extract meaningful biological insights.
Purpose of the Study:
- To develop and demonstrate a modular analytics pipeline for unifying heterogeneous RNA-seq datasets.
- To enable robust biomarker discovery by integrating public data using open-source tools and machine learning.
- To showcase the pipeline's adaptability across different disease contexts and data types.
Main Methods:
- Developed a modular pipeline incorporating quality control, differential expression, and pathway analysis.
- Integrated competitive machine learning techniques to unify disparate RNA-seq datasets.
- Applied the pipeline to COVID-19 severity, multi-cohort sepsis, and atherosclerosis (bulk and single-cell) datasets.
Main Results:
- Successfully identified differentially expressed gene signatures for COVID-19 severity.
- Generated concise biomarker panels for sepsis using integrated differential expression and machine learning.
- Examined tissue and cell-type specificity of N-acyl-phosphatidylethanolamine phospholipase D in atherosclerosis.
Conclusions:
- An adaptable, modular pipeline using open-source tools can effectively repurpose public RNA-seq data.
- The pipeline reduces noise, generates novel hypotheses, and reveals biological insights for biomarker discovery.
- This approach establishes a foundation for future research leveraging public data.
