Related Experiment Video
Updated: Jul 9, 2026

Modeling an Enzyme Active Site using Molecular Visualization Freeware
Published on: December 25, 2021
Deciphering key factors of active learning performance in biomolecular design
Yixuan Zhi1, Qixiu Du1, Han Yu1
1Ministry of Education Key Laboratory of Bioinformatics, Center for Synthetic and Systems Biology, Beijing National Research Center for Information Science and Technology, Department of Automation, Tsinghua University, Beijing 100084, China.
Active learning (AL) improves biomolecule design by optimizing fitness. Optimization efficiency depends on initial settings and data properties, not just the AL algorithm, guiding better strategy selection.
Area of Science:
- Biomolecular engineering
- Computational biology
- Machine learning
Background:
- Machine learning (ML) and active learning (AL) are increasingly used for efficient biomolecule design.
- Challenges in genetic engineering include complex, noisy, and sparse sequence-fitness relationships, hindering reliable AL evaluation.
- A comprehensive benchmark is needed to assess AL performance determinants in biomolecular design.
Purpose of the Study:
- To benchmark active learning (AL) algorithms for biomolecular sequence design.
- To identify key factors influencing AL efficiency in optimizing sequence fitness.
- To propose metrics for selecting optimal AL strategies in genetic engineering.
Main Methods:
- Developed a benchmark using large-scale protein and DNA regulatory sequence libraries.
- Evaluated uncertainty quantification (UQ) algorithms for calibration and accuracy.
- Assessed various sampling strategies and factors affecting optimization efficiency.
Main Results:
- Ensemble-based UQ algorithms demonstrated robustness and generality.
- No single sampling strategy was universally optimal; greedy iterative strategies showed practical utility.
- Optimization efficiency was primarily influenced by initial settings, distribution sparsity, and sequence similarity, not the specific AL algorithm.
Conclusions:
- Biomolecular sequence design efficiency using AL is more dependent on data characteristics and initial setup than the chosen AL algorithm.
- Proposed quantifiable metrics can aid in selecting appropriate AL strategies for practical applications.
- Findings provide valuable insights for implementing AL pipelines in genetic engineering.
More Related Videos
Related Concept Videos
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence its...
Biopharmaceutical Factors Influencing Drug Product Design: Overview
Factors Affecting Activity Coefficient
The activity coefficient value for an ion is close to one when the solution has almost zero ionic strength, i.e., when the solution shows close to ideal behavior. As the ionic strength of the solution increases from 0 to 0.1 mol/L, a decrease in the...
Biopharmaceutics and Pharmacokinetics: Overview
Noncovalent Attractions in Biomolecules
Four types of noncovalent interactions are hydrogen bonds, van der Waals forces, ionic bonds, and hydrophobic interactions.
Hydrogen bonding results from the electrostatic attraction of a hydrogen atom covalently bonded to a strong-electronegative atom like oxygen,...
Introduction to Mechanisms of Enzyme Catalysis

