Experimental and analytical pipeline for sub-genomic RNA landscape of coronavirus by Nanopore sequencer

Bo-Jia Chen1, Ching-Hung Lin2, Hung-Yi Wu2

  • 1Doctoral Program in Microbial Genomics, National Chung Hsing University and Academia Sinica, Taichung, Taiwan.

Microbiology Spectrum
|March 14, 2024
PubMed

Insights

This study establishes optimized bioinformatics parameters for Nanopore sequencing of coronavirus defective viral genomes (DVGs). It provides crucial benchmarks for accurate viral RNA transcriptome analysis, improving DVG identification in bovine coronavirus (BCoV) research.

Area of Science:

  • Virology
  • Bioinformatics
  • Molecular Biology

Background:

  • Coronaviruses (CoVs) pose significant medical and economic threats, with their life cycle involving complex subgenomic RNAs, including defective viral genomes (DVGs).
  • Previous studies using Nanopore sequencing for viral RNA transcriptomes lacked standardized bioinformatics parameters, leading to unreliable results.
  • Investigating bovine coronavirus (BCoV) offers a safer alternative (biosafety level 2) for validating sequencing and bioinformatics approaches.

Purpose of the Study:

  • To establish optimized and validated bioinformatics parameters for Nanopore sequencing of coronavirus subgenomic RNA transcriptomes.
  • To rigorously assess and refine bioinformatic pipelines for accurate identification of defective viral genomes (DVGs).
  • To provide benchmarks for library construction and data analysis in coronavirus RNA research.

Main Methods:

  • Utilized bovine coronavirus (BCoV) for experiments under biosafety level 2 conditions.
  • Employed four Nanopore (ONT) sequencing protocols (RNA direct, cDNA direct, with/without exonuclease treatment).
  • Rigorously optimized bioinformatics parameters (k-mer, gap, segment, bin sizes) and validated findings using traditional PCR and Sanger sequencing.

Main Results:

  • Developed an optimized bioinformatics pipeline achieving 82.6% sensitivity and 99.6% specificity for DVG read identification.
  • Determined optimal cutoffs for bioinformatic parameters to effectively remove sequence noise while retaining informative DVG reads.
  • Exonuclease treatment was found to reduce RNA transcript abundance but was not essential for library preparation.

Conclusions:

  • This study provides essential benchmarks for library preparation and bioinformatics analysis of the discontinuous coronavirus RNA transcriptome.
  • The optimized pipeline enhances the accuracy of identifying viral RNA species, including functional DVGs.
  • Further investigation into the biological functions of identified BCoV RNA sequences is warranted.