Related Experiment Video
Updated: Aug 6, 2026

Using RNA-sequencing to Detect Novel Splice Variants Related to Drug Resistance in In Vitro Cancer Models
Published on: December 9, 2016
BlendSplice: A Frequency-Blended Generative Framework for In Silico Synthesis of Biologically Realistic Splice Site
Espoir Kabanga1,2, Seonil Jee2, Arnout Van Messem3
1IDLab, Department of Electronics and Information Systems, Ghent University, Ghent 9000, Belgium.
None:
Generative models for biological sequences face challenges in balancing sequence realism with diversity. We investigated whether posttraining frequency blending, combining model-learned distributions with empirical nucleotide priors, can improve synthetic-sequence quality across diverse generative architectures. We present BlendSplice, a frequency-blended generative framework for the in silico synthesis of biologically realistic splice site sequences. Our approach combines position-specific empirical nucleotide priors with 3 generative architectures: a one-dimensional U-Net-based denoising diffusion probabilistic model, a generative adversarial network, and a variational autoencoder. By linearly blending model-learned distributions with empirical base frequencies through a tunable weight parameter lambda ( ), we enable explicit control over the realism-diversity trade-off. We evaluate synthetic donor (5') and acceptor (3') splice sites from Arabidopsis thaliana, Homo sapiens, and Danio rerio, at 2 sequence lengths (402 and 2,002 bp), at (no blending) and (balanced blending of priors and model probabilities) using direct assessments (GC content, nucleotide conservation, sequence logos, -mer divergence, and polypyrimidine tract analysis) and indirect functional validation via state-of-the-art splice site classifiers (SpliceRover and Spliceator) and independent validation with SpliceAI, under 3 evaluation protocols: Train-Real/Test-Synthetic (biological realism), Train-Synthetic/Test-Real (transferability), and fixed-budget data augmentation with varying real/synthetic mixtures. Our results demonstrate that frequency blending at substantially improves biological fidelity and downstream predictive performance for the generative adversarial network and diffusion models, whereas the variational autoencoder model generates high-quality sequences even without blending. Data augmentation with a 50% real and 50% synthetic mix achieves predictive performance comparable to the baseline across all 3 species, indicating that blending is an effective strategy for generating high-quality synthetic genomic data.
Related Concept Videos
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
RNA Splicing
RNA Splicing
Lagging Strand Synthesis
There are several major differences between synthesis of the leading strand and synthesis of the lagging strand. 1) Leading strand synthesis happens in the direction of replication fork opening, whereas lagging strand synthesis happens in the...
Lagging Strand Synthesis
There are several major differences between synthesis of the leading strand and synthesis of the lagging strand. 1) Leading strand synthesis happens in the direction of replication fork opening, whereas lagging strand synthesis happens in the...
Alternative RNA Splicing
There are five types of alternative RNA splicing that vary in the ways the pre-mRNA segments are removed or retained in the mature mRNA. The first...
