Identification of Cancer Genes Based on De Novo Transposon Insertion Site Analysis Using RNA and DNA Sequencing

Aaron Sarver1,2

  • 1Institute for Health Informatics, University of Minnesota, Minneapolis, MN, USA. sarver@umn.edu.

Insights

A new bioinformatics pipeline identifies cancer-causing transposon insertion sites in mouse genomes. This method aids in discovering novel cancer genes by analyzing both DNA and RNA sequencing data.

Area of Science:

  • Genetics
  • Bioinformatics
  • Cancer Research

Background:

  • Forward genetic screens using insertional mutagenesis are crucial for identifying cancer genes.
  • The Sleeping Beauty DNA transposon is widely used to induce mutations and cause cancer in murine models.
  • Accurate identification of transposon insertion sites is essential for cancer gene discovery.

Purpose of the Study:

  • To present a versatile bioinformatics pipeline for identifying transposon insertion sites.
  • To enable the detection of transposon-generated gene fusions in RNA sequencing data.
  • To facilitate direct identification of transposon insertion sites in DNA sequencing data.

Main Methods:

  • Development of a bioinformatics pipeline for analyzing sequencing data.
  • Application of the pipeline to RNA sequencing data for fusion identification.
  • Application of the pipeline to DNA sequencing data for direct insertion site identification.

Main Results:

  • The pipeline successfully identifies transposon-generated fusions in RNA-Seq data.
  • The pipeline accurately detects direct transposon insertion sites in DNA sequencing data.
  • The method is currently used to analyze Sleeping Beauty transposon insertions in the murine genome.

Conclusions:

  • The presented bioinformatics pipeline is effective for identifying transposon insertions.
  • This approach aids in the discovery of candidate cancer genes.
  • The method is adaptable for identifying any mobile genetic element in various genomes.

Related Concept Videos

DNA-only Transposons02:57

DNA-only Transposons

DNA-only transposons are called autonomous transposons since they code for the enzyme transposase that is required for the transposition mechanism. Insertion of transposons can alter gene functions in multiple ways. They can mutate the gene, alter gene expression by introducing a novel promoter or insulator sequence, introduce new splice sites, and change the mRNA transcripts produced, or remodel chromatin structure.
The donor site from where the transposon is excised is either degraded or...
17.4K
Bacterial RNA Polymerase00:43

Bacterial RNA Polymerase

Unlike eukaryotes, bacteria use a single RNA Polymerase (RNAP) to transcribe all genes. The different subunits of bacterial RNAPhave distinct functions. The multisubunit structure of the bacterial RNAP helps the enzyme to maintain catalytic function, facilitate assembly, interact with DNA and RNA, and self-regulate its activity.
In most genes, the transcription site is a single base present upstream of the coding sequence. Though RNAP is a catalytically efficient enzyme, it does not recognize...
32.8K
From DNA to Protein03:06

From DNA to Protein

The flow of genetic information in cells from DNA to mRNA to protein is described by the central dogma, which states that genes specify the sequence of mRNAs, which in turn specify the sequence of amino acids making up all proteins. The decoding of one molecule to another is performed by specific proteins and RNAs. Because the information stored in DNA is so central to cellular function, it makes intuitive sense that the cell would make mRNA copies of this information for protein synthesis...
22.4K
DNA Base Pairing02:27

DNA Base Pairing

Erwin Chargaff’s rules on DNA equivalence paved the way for the discovery of base pairing in DNA. Chargaff’s rules state that in a double-stranded DNA molecule,
33.2K
Eukaryotic RNA Polymerases00:58

Eukaryotic RNA Polymerases

RNA Polymerase (RNAP) is conserved in all animals, with bacterial, archaeal, and eukaryotic RNAPs sharing significant sequence, structural, and functional similarities. Among the three eukaryotic RNAPs, RNA Polymerase II is most similar to bacterial RNAP in terms of both structural organization and folding topologies of the enzyme subunits. However, these similarities are not reflected in their mechanism of action.
All three eukaryotic RNAPs require specific transcription factors, of which the...
27.1K
RNA Splicing01:32

RNA Splicing

Splicing is the process by which eukaryotic RNA is edited before its translation into protein. The RNA strand transcribed from eukaryotic DNA is called the primary transcript. The primary transcripts that become mRNAs are called precursor messenger RNAs (pre-mRNAs). Eukaryotic pre-mRNA contains alternating sequences of exons and introns. Exons are nucleotide sequences that code for proteins, whereas introns are the non-coding regions. In RNA splicing, introns are removed and exons are bonded...
60.6K