Related Experiment Video
Updated: Mar 18, 2026

Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens
Published on: June 28, 2018
Unlocking Enzyme Discovery: A High-Resolution Gene Cluster Database Powered by Phylogenetic Insights and Machine
Sidun Zhang1, Junlong He1, Xuguo Duan2
1School of Biotechnology, Jiangnan University, 1800 Lihu Road, Wuxi 214122, China.
Abstract:
High-throughput sequencing has generated vast genomic repositories that remain under-annotated, hampering enzyme discovery. We present an integrated pipeline that (i) builds a high-resolution, cross-kingdom phylogenetic database, (ii) mines candidates via multilocus phylogeny, (iii) predicts activities using an evolutionary-scale protein language model, and (iv) removes false positives through multilevel residue-atom contact rescoring. When applied to the r-BOX pathway, this approach uncovered numerous previously undocumented FadB, BktB, Ter, and YdiI homologues. Our activity model achieved R2 = 0.68 and reduced the RMSE on high-value targets by 11% compared to the prior SOTA (UniKP). Contact scoring improved early enrichment (EF1%) by 16-fold. Experimental validation targeting FadB increased titers from 0.65 g/L (shake flasks) to 1.7 g/L, reaching 10.2 g/L in a fermentation process. Together, these results establish a robust, generalizable framework for discovering scarce, high-value enzymes and prioritizing functional variants at scale.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Gene Evolution - Fast or Slow?
Phylogeny
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...

