Related Experiment Video
Updated: May 7, 2026

High-throughput Screening for Chemical Modulators of Post-transcriptionally Regulated Genes
Published on: March 3, 2015
Bridging the gap between transcriptome and proteome measurements identifies post-translationally regulated genes
Yawwani Gunawardana1, Mahesan Niranjan
1School of Electronics and Computer Science, University of Southampton, Southampton, SO17 1BJ, UK.
This study developed a machine learning model to predict protein concentrations from messenger RNA (mRNA) levels in Saccharomyces cerevisiae. Outliers in prediction suggest post-translational regulation, revealing novel regulatory insights.
Area of Science:
- Molecular Biology
- Systems Biology
- Computational Biology
Background:
- Messenger RNA (mRNA) abundance is often used as a proxy for protein levels, but this correlation is not universal due to post-transcriptional and post-translational regulation.
- High-throughput omic measurements provide valuable data, but bridging the gap between transcriptome and proteome levels remains a challenge.
Purpose of the Study:
- To develop a data-driven machine learning approach to predict protein concentrations from mRNA data in Saccharomyces cerevisiae.
- To identify mRNA-protein pairs potentially regulated by post-translational modifications.
Main Methods:
- Applied feature selection using sparsity-inducing regression (l1 norm regularization) to identify key predictive features.
- Reduced 37 features to 5: mRNA, ribosomal occupancy, ribosome density, tRNA adaptation index, and codon bias.
- Developed a linear predictor to model the relationship between selected features and protein concentrations.
Main Results:
- Achieved accurate prediction of protein concentrations with R² = 0.86 using the reduced feature set.
- Identified proteins with inaccurately predicted concentrations (outliers) as candidates for post-translational regulation.
- Demonstrated significantly higher evidence of post-translational modification in outliers compared to random subsets (P < 0.02).
Conclusions:
- Machine learning models can effectively bridge the transcriptome-proteome gap.
- Analysis of prediction outliers offers a novel method to uncover post-translational regulatory mechanisms.
- Outliers in machine learning models can yield significant biological insights.
Related Concept Videos
Ribosome Profiling
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...
Regulation of Expression at Multiple Steps
Regulation of Expression Occurs at Multiple Steps
Transcription results in the generation of precursor (pre-mRNA) that consists of both exons and introns, which needs further processing before being translated to a...
What is Gene Expression?
What is Gene Expression?
Gene expression is the process in which DNA directs the synthesis of functional products, that is, proteins. Cells can regulate gene expression at various stages. It allows organisms to generate different cell types and enables cells to adapt to internal and external factors.
Genetic Information Flows from DNA to RNA to Protein
A gene is a stretch of DNA that serves as the blueprint for functional RNAs and proteins. Since DNA is made up of nucleotides and proteins consist of amino...
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...

