Related Experiment Video
Updated: Jul 9, 2025

07:09
A Bioinformatics Pipeline for Investigating Molecular Evolution and Gene Expression using RNA-seq
Published on: May 28, 2021
9.6K
Software pipelines for RNA-Seq, ChIP-Seq and germline variant calling analyses in common workflow language (CWL).
Konstantinos A Kyritsis1, Nikolaos Pechlivanis1,2, Fotis Psomopoulos1
1Institute of Applied Biosciences (INAB), Centre for Research and Technology Hellas (CERTH), Thessaloniki, Greece.
Frontiers in Bioinformatics
|November 29, 2023
Summary
Automated bioinformatics pipelines using Common Workflow Language (CWL) ensure reproducible analysis of high-throughput sequencing (HTS) data. These CWL workflows accurately reproduce results and detect genetic variants, facilitating research in areas like Chronic Lymphocytic Leukemia (CLL).
Area of Science:
- Bioinformatics and Computational Biology
- Genomics and Molecular Biology
- Data Science and Machine Learning
Background:
- Reproducibility in data analysis is crucial, especially for large-scale High-throughput Sequencing (HTS) datasets.
- Automating data analysis pipelines addresses the need for consistent and reliable results in genomics research.
- Challenges in bioinformatics include software incompatibility and complex configuration, hindering reproducibility.
Purpose of the Study:
- To develop and evaluate automated Common Workflow Language (CWL) pipelines for HTS data analysis.
- To assess the performance of CWL workflows in reproducing published results and detecting genetic variants.
- To provide a flexible, reusable, and open-source resource for standard bioinformatics analyses.
Main Methods:
- Implementation of automated workflows using Common Workflow Language (CWL).
- Analysis of RNA-Seq, ChIP-Seq, and Germline variant calling experiments.
- Validation through reproduction of Chronic Lymphocytic Leukemia (CLL) study results and analysis of Genome in a Bottle (GIAB) Consortium data.
Main Results:
- CWL-implemented workflows demonstrated high accuracy in reproducing prior findings.
- Successful identification of significant biomarkers in Chronic Lymphocytic Leukemia (CLL) datasets.
- Accurate detection of germline Single Nucleotide Polymorphism (SNP) and small Insertion/Deletion (INDEL) variants in whole genome sequencing data.
Conclusions:
- CWL pipelines offer significant advantages in reproducibility and reusability for bioinformatics analyses.
- Containerization combined with CWL overcomes software compatibility and configuration challenges.
- The developed CWL workflows are flexible, adaptable, and serve as an open resource for short-read data analysis, promoting automation and cross-platform compatibility.

