Related Experiment Video
Updated: Jun 20, 2025

Author Spotlight: Exploring Cellular Processes by Modeling Ligands in Cryo-EM Maps
Published on: July 19, 2024
Identifying and embedding transferability in data-driven representations of chemical space
Tim Gould1, Bun Chan2, Stephen G Dale1,3
1Queensland Micro- and Nanotechnology Centre, Griffith University Nathan Qld 4111 Australia.
We developed a transferability assessment tool (TAT) to identify and embed transferable data for machine learning in chemistry. This method overcomes human biases in data curation, improving model generalization for density functional approximations (DFAs).
Area of Science:
- Computational Chemistry
- Materials Science
- Machine Learning
Background:
- Transferability is crucial for model generalization across scientific disciplines.
- Complex machine learning models can obscure how data transferability is achieved or missed.
- Identifying and embedding transferable training data are key challenges in data-driven scientific modeling.
Purpose of the Study:
- To develop methods for identifying and embedding transferable training data for ab initio chemical modeling.
- To address the challenge of understanding and improving data transferability in complex machine learning models.
- To enhance the generalization capabilities of data-driven models in chemistry.
Main Methods:
- Introduction of a Transferability Assessment Tool (TAT) for ab initio chemical modeling.
- Demonstration of TAT on a controllable data-driven model for developing density functional approximations (DFAs).
- Analysis of human intuition's impact on training data curation and its chemical biases.
Main Results:
- Human intuition in data curation can introduce chemical biases, hindering the transferability of data-driven DFAs.
- The TAT successfully identified key factors influencing data transferability.
- Three transferability principles were motivated, including the concept of transferable diversity.
Conclusions:
- The developed TAT and identified transferability principles offer strategies to improve data curation for machine learning in chemistry.
- Embedding transferable diversity into training data can enhance the generalization of general-purpose machine learning models.
- This work provides a framework for creating more robust and transferable data-driven models in scientific research.
Related Concept Videos
Molecular Models
Inductive Effects on Chemical Shift: Overview
Chemical Shift: Internal References and Solvent Effects
The internal reference compound generally used in NMR spectroscopy is tetramethylsilane (TMS). TMS is preferred because it is chemically inert, soluble in NMR solvents, and easily removable. Also, the highly shielded methyl protons in TMS yield an intense...
¹H NMR Chemical Shift Equivalence: Homotopic and Heterotopic Protons
Fischer Projections
Resonance and Hybrid Structures
Resonance Structures and Resonance Hybrids
The Lewis structure of a nitrite anion (NO2−) may actually be drawn in two different ways, distinguished by the locations of the N–O and N=O bonds.

