Related Experiment Video
Updated: May 25, 2026

10:49
Measuring RAN Peptide Toxicity in C. elegans
Published on: April 30, 2020
Polyglutamine repeats are associated to specific sequence biases that are conserved among eukaryotes
Matteo Ramazzotti1, Elodie Monsellier, Choumouss Kamoun
1Dipartimento di Scienze Biochimiche, Università degli Studi di Firenze, Florence, Italy. matteo.ramazzotti@unifi.it
Plos One
|February 8, 2012
Summary
Protein aggregation linked to neurodegenerative diseases may be prevented by specific amino acid sequences. These sequences, found in eukaryotic proteomes, suggest evolutionary selection against harmful polyglutamine expansions.
Area of Science:
- Molecular Biology
- Genetics
- Neuroscience
Background:
- Nine human neurodegenerative diseases are linked to protein aggregation caused by expanded polyglutamine (polyQ) tracts.
- PolyQ expansion is thought to result from polyCAG codon expansion during replication.
- However, many polyQ proteins remain soluble, suggesting counter-selection mechanisms.
Purpose of the Study:
- To identify the genetic and protein contexts that may counteract polyQ expansion and aggregation.
- To investigate evolutionary pressures shaping polyQ-containing proteins.
Main Methods:
- Developed custom software to analyze entire eukaryotic proteomes for imperfect polyQ sequences.
- Assessed amino acid residues flanking polyQ tracts and non-glutamine residues within polyQ sequences.
- Examined 15 eukaryotic proteomes for conserved amino acid biases associated with polyQ regions.
Main Results:
- Discovered significant amino acid biases associated with polyQ regions across 15 eukaryotic proteomes.
- Observed over-representation of Proline, Leucine, and Histidine, and under-representation of Aspartic acid, Cysteine, and Glycine.
- These biases are conserved across unrelated proteins and independent of protein function.
Conclusions:
- Specific amino acid residues appear to have been co-selected with polyQ sequences during evolution.
- These findings suggest a protective role for certain amino acids against polyQ expansion and aggregation.
- Further research is needed to elucidate the selective pressures driving these observed amino acid biases.
More Related Videos
Related Concept Videos
Multi-species Conserved Sequences
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
From DNA to Protein
The flow of genetic information in cells from DNA to mRNA to protein is described by the central dogma, which states that genes specify the sequence of mRNAs, which in turn specify the sequence of amino acids making up all proteins. The decoding of one molecule to another is performed by specific proteins and RNAs. Because the information stored in DNA is so central to cellular function, it makes intuitive sense that the cell would make mRNA copies of this information for protein synthesis...
Leaky Scanning
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R stands for...
Signal Sequences and Sorting Receptors
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...

