Quantifying the Hardness of Bioactivity Prediction Tasks for Transfer Learning.
Hosein Fooladi1,2,3, Steffen Hirte1,3, Johannes Kirchmair1,2
1Department of Pharmaceutical Sciences, Division of Pharmaceutical Chemistry, Faculty of Life Sciences, University of Vienna, Josef-Holaubek-Platz 2, 1090 Vienna, Austria.
Journal of Chemical Information and Modeling
|May 13, 2024
Summary
This study introduces a novel method to predict the difficulty of bioactivity prediction tasks in machine learning for drug discovery. Quantifying task hardness helps estimate performance gains from knowledge-sharing methods like meta-learning.
Area of Science:
- Computational chemistry
- Machine learning in drug discovery
- Bioactivity prediction
Background:
- Machine learning (ML) is crucial in drug discovery but hindered by data scarcity.
- Knowledge-sharing ML strategies like transfer learning, multitask learning, and meta-learning address data limitations by leveraging related tasks.
- A key challenge is understanding how source task relatedness impacts the performance of these ML models.
Purpose of the Study:
- To develop a new method for quantifying and predicting the hardness of bioactivity prediction tasks.
- To assess the relationship between task hardness and the performance of knowledge-sharing ML approaches.
Main Methods:
- Generated protein and chemical representations.
- Calculated distances between the target bioactivity prediction task and available training tasks to quantify task hardness.
- Applied the metric within a meta-learning framework on the FS-Mol dataset.
Main Results:
- Demonstrated an inverse correlation between the proposed task hardness metric and model performance (Pearson's r = -0.72) in a meta-learning scenario.
- The developed metric effectively quantifies the difficulty of a bioactivity prediction task relative to existing training data.
Conclusions:
- The new task hardness metric is a valuable tool for estimating performance improvements achievable with meta-learning in drug discovery.
- This metric can guide the application and development of knowledge-sharing ML strategies to overcome data scarcity challenges.
Related Concept Videos
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Toughness and Hardness of Aggregate
Toughness and hardness are critical properties of aggregate materials used in concrete, particularly on pavement surfaces and industrial flooring subjected to heavy loads. Toughness is defined as the aggregate's resistance to failure by impact and is measured by the aggregate impact value (AIV). For this, the aggregate impact value test is performed, wherein the impact is delivered by a standard hammer, which falls freely under its own weight onto the aggregates. The aggregates fragment in the...


