Related Experiment Video
Updated: Sep 7, 2025

Annotation of Plant Gene Function via Combined Genomics, Metabolomics and Informatics
Published on: June 17, 2012
Predicting Monoterpene Indole Alkaloid-Related Genes from Expression Data with Artificial Neural Networks
Thomas Dugé de Bernonville1, Emily Amor Stander2, Géraud Dugé de Bernonville3
1EA2106 Biomolécules et Biotechnologies Végétales, Université de Tours, Tours, France. thomas.dugedebernonville@limagrain.com.
Abstract:
Elucidation of biological pathways leading to specialized metabolites remains a complex task. It is however a mandatory step to allow bioproduction into heterologous hosts. Many steps have already been identified using conventional approaches, enlarging the space of known possible chemical steps. In the recent past years, identification of missing steps has been fueled by the generation of genomic and transcriptomic data for nonmodel species. The analysis of gene expression profiles has revealed that in many cases, genes encoding enzymes involved in the same biosynthetic pathways are coexpressed across different tissue types and environmental conditions. Hence, coexpressed studies, either in the form of differential gene expression, gene coexpression network, or unsupervised clustering methods, have helped deciphering missing steps to complete knowledge on biosynthetic pathways. Already identified biosynthetic steps can be used as baits to capture the remaining unknown steps. The present protocol shows how supervised machine learning in the form of artificial neural networks (ANNs) can efficiently classify genes as specialized metabolism related or not according to their expression levels. Using Catharanthus roseus as an example, we show that ANN trained on a minimal set of bait genes results in many true positives (correctly predicted genes) while keeping false positives low (containing possible candidate genes).

