Related Experiment Video
Updated: Oct 7, 2026

Cost-Efficient Transcriptomic-Based Drug Screening
Published on: February 23, 2024
RandomReadsMG: Rapid, realistic metagenome simulation at terabase scale for benchmarking and experimental design
Brian Bushnell1, Frederik Schulz1, Juan C Villada1
1DOE Joint Genome Institute, Lawrence Berkeley National Laboratory, Berkeley, CA 94720, USA.
Abstract:
Realistic synthetic metagenomes are essential for benchmarking bioinformatics tools, guiding experimental design, and generating labeled data for artificial intelligence and machine learning. We present RandomReadsMG, an open-source metagenomic read simulator distributed with BBTools. It generates communities containing hundreds to 50,000 genomes in a single command and supports user-defined abundance profiles, configurable within-genome coverage variation, retained read provenance, sequencing errors, and library artifacts. Presets are provided for Illumina, Oxford Nanopore Technologies (ONT), and PacBio sequencing. In benchmarks, RandomReadsMG produced terabase-scale datasets in under 6 h while maintaining bounded memory use. Simulations based on empirical drinking water profiles preserved genome-level depth after remapping. Controlled pathogen spike-ins further showed how simulation can estimate thresholds for read detection and metagenome-assembled genome (MAG) recovery. RandomReadsMG provides a fast, reproducible framework for metagenomics benchmarking, biosurveillance study design, and large-scale synthetic data generation.

