Related Experiment Video
Updated: Jun 30, 2026

CIRCLE-Seq for Interrogation of Off-Target Gene Editing
Published on: November 1, 2024
Poisoning the Genome: Targeted Backdoor Attacks on DNA Foundation Models
Charalampos Koilakos1, Ioannis Mouratidis1, Ilias Georgakopoulos-Soares1
1Division of Pharmacology and Toxicology, College of Pharmacy, The University of Texas at Austin, Dell Pediatric Research Institute, Austin, TX, USA.
Genomic foundation models are vulnerable to data poisoning attacks, where even a small percentage of malicious data can degrade performance on specific biological tasks. This highlights the need for robust data integrity checks in AI model development.
Area of Science:
- Genomics
- Artificial Intelligence
- Bioinformatics
Background:
- Foundation models trained on DNA sequences excel at biological tasks like variant effect prediction.
- These models utilize massive genomic datasets, but DNA's lack of semantic transparency hinders detection of corrupted data.
- Genomic data curation faces challenges in identifying adversarial entries.
Purpose of the Study:
- To systematically investigate data poisoning vulnerabilities in genomic language models.
- To assess the impact of poisoning during both pre-training and fine-tuning stages.
- To evaluate the susceptibility of models to targeted attacks on specific genomic features and tasks.
Main Methods:
- Investigated data poisoning in Evo 2 and GENERator architectures during pre-training.
- Simulated attacks by corrupting TATA-box motifs, CTCF binding sites, and inserting synthetic sequences.
- Explored fine-tuning attacks including backdoor installation via CTCF site poisoning and label corruption for variant classification.
Main Results:
- Less than 1% poisoned data at pre-training selectively degraded generative performance on targeted genomic contexts.
- Fine-tuning attacks successfully installed conditional backdoors and compromised variant classification tasks (e.g., BRCA1).
- Genomic foundation models demonstrated susceptibility to targeted data poisoning with minimal footprint.
Conclusions:
- Genomic foundation models are vulnerable to sophisticated data poisoning attacks.
- The findings underscore the critical need for enhanced data security and validation in AI for genomics.
- Recommended adopting data provenance tracking, integrity verification, and adversarial robustness evaluation.
Related Concept Videos
Overview of DNA Repair
Chemically...
Overview of DNA Repair
Chemically...
DNA Damage can Stall the Cell Cycle
DNA Damage Can Stall the Cell Cycle
Genome Copying Errors
Base Excision Repair
The first step of...
