A Bayesian model based computational analysis of the relationship between bisulfite accessible single-stranded DNA in

Guojun Yu1, Yingru Wu2, Zhi Duan1

  • 1Department of Cell Biology, Albert Einstein College of Medicine, Bronx, New York, United States of America.

Plos Computational Biology
|September 7, 2021
PubMed

The B cells in our body generate protective antibodies by introducing somatic hypermutations (SHM) into the variable region of immunoglobulin genes (IgVs). The mutations are generated by activation induced deaminase (AID) that converts cytosine to uracil in single stranded DNA (ssDNA) generated during transcription. Attempts have been made to correlate SHM with ssDNA using bisulfite to chemically convert cytosines that are accessible in the intact chromatin of mutating B cells. These studies have been complicated by using different definitions of "bisulfite accessible regions" (BARs). Recently, deep-sequencing has provided much larger datasets of such regions but computational methods are needed to enable this analysis. Here we leveraged the deep-sequencing approach with unique molecular identifiers and developed a novel Hidden Markov Model based Bayesian Segmentation algorithm to characterize the ssDNA regions in the IGHV4-34 gene of the human Ramos B cell line. Combining hierarchical clustering and our new Bayesian model, we identified recurrent BARs in certain subregions of both top and bottom strands of this gene. Using this new system, the average size of BARs is about 15 bp. We also identified potential G-quadruplex DNA structures in this gene and found that the BARs co-locate with G-quadruplex structures in the opposite strand. Using various correlation analyses, there is not a direct site-to-site relationship between the bisulfite accessible ssDNA and all sites of SHM but most of the highly AID mutated sites are within 15 bp of a BAR. In summary, we developed a novel platform to study single stranded DNA in chromatin at a base pair resolution that reveals potential relationships among BARs, SHM and G-quadruplexes. This platform could be applied to genome wide studies in the future.

Related Concept Videos

Mismatch Repair01:20

Mismatch Repair

Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
5.4K
Conservative Site-specific Recombination and Phase Variation02:53

Conservative Site-specific Recombination and Phase Variation

Because the DNA segments are cut and reorganized in a direction-specific manner, site-specific recombination has emerged as an efficient genetic engineering technique. Flippase and Cyclization recombinases or Flp and Cre, respectively, are two members of the tyrosine recombinase family derived from bacteriophages, that are used to mediate site-specific DNA insertions, deletions, and targeted expression of proteins in mammalian cell lines.
The recognition sites for Cre recombinase called LoxP...
6.3K
Crossing Over01:30

Crossing Over

Crossing over is the exchange of genetic information between homologous chromosomes during prophase I of meiosis I. Genetic recombination gives rise to allelic diversity in the newly formed daughter cells. In humans, crossing over produces genetically distinct haploid egg and sperm cells that undergo fertilization to produce unique offspring. Before cell division starts, the germ cell’s chromosome(s) undergo duplication in the S phase of the cell cycle. As the cells enter prophase I,...
5.0K
Gene Conversion02:08

Gene Conversion

Other than maintaining genome stability via DNA repair, homologous recombination plays an important role in diversifying the genome. In fact, the recombination of sequences forms the molecular basis of genomic evolution. Random and non-random permutations of genomic sequences create a library of new amalgamated sequences. These newly formed genomes can determine the fitness and survival of cells. In bacteria, homologous and non-homologous types of recombination lead to the evolution of new...
10.1K
Spreading of Chromatin Modifications02:25

Spreading of Chromatin Modifications

The histone proteins in the nucleosomes are post-translationally modified (PTM) to increase or decrease access to DNA. The commonly observed PTMs are methylation, acetylation, phosphorylation, and ubiquitination of lysine amino acids in the histone H3 tail region. These histone modifications have specific meaning for the cell. Hence, they are called "histone code". The protein complex involved in histone modification is termed as "reader-writer" complex.
Writers
The writer...
8.8K