Related Experiment Video
Updated: Jul 9, 2025

Author Spotlight: In Silico Creation and Impact of Carbonylated Amino Acids on Protein Structure and Function
Published on: April 26, 2024
Transferring a Molecular Foundation Model for Polymer Property Predictions.
Pei Zhang1, Logan Kearney2, Debsindhu Bhowmik1
1Computational Sciences and Engineering Division, Oak Ridge National Laboratory, Oak Ridge, Tennessee 37831, United States.
Large language models (LLMs) can accelerate scientific discovery. By pretraining LLMs on small molecules and fine-tuning on polymer data, researchers achieved accurate polymer property predictions without costly data augmentation.
Area of Science:
- Computational chemistry
- Materials science
- Polymer science
Background:
- Transformer-based large language models (LLMs) show promise for accelerating design optimization in fields like drug development and material discovery.
- Self-supervised pretraining of LLMs necessitates large datasets, which are scarce in specialized domains such as polymer science.
- Current methods for polymers involve data augmentation, increasing computational expenses.
Purpose of the Study:
- To investigate the efficacy of transfer learning for polymer property prediction using LLMs.
- To address data scarcity in polymer science by leveraging readily available small molecule datasets.
- To compare the performance of transfer learning against traditional data augmentation techniques.
Main Methods:
- Utilized transformer-based LLMs pretrained on extensive small molecule datasets.
- Fine-tuned the pretrained models on specific polymer property datasets.
- Evaluated model accuracy on benchmark prediction tasks for polymers.
- Compared results against models trained solely on augmented polymer data.
Main Results:
- LLMs pretrained on small molecules and fine-tuned on polymers achieved accuracy comparable to models trained on augmented polymer data.
- Transfer learning proved effective in overcoming data scarcity in polymer science.
- This approach offers a computationally efficient alternative to data augmentation.
Conclusions:
- Transfer learning with LLMs offers a viable and efficient strategy for polymer property prediction.
- Pretraining on small molecules provides a robust foundation for models applied to polymer science.
- This method reduces the need for extensive data augmentation, saving computational resources.
More Related Videos
Related Concept Videos
Polymers: Defining Molecular Weight
The number average molecular weight (Mn) is the summation of the number...
Polymers: Molecular Weight Distribution
Molecular Weight of Step-Growth Polymers
As the step-growth polymerization involves step-wise condensation of monomers, the molecular weight also builds up eventually. Consequently, high molecular weight polymers are obtained at the late stages of the polymerization, where 99% of monomers have been consumed.
The extent of the...
Step-Growth Polymerization: Overview
Many natural and synthetic polymers are produced by...
Polymer Classification: Crystallinity
Crystalline domains are the regions where polymer chains are aligned in an orderly manner and held together in proximity by intermolecular forces. For example, chains in the crystalline domains of polyethylene and nylon are bound together by van der Waals...
Molecular Models

