Related Experiment Video
Updated: May 28, 2025

DNA Nanotubes as a Versatile Tool to Study Semiflexible Polymers
Published on: October 25, 2017
Harnessing large language models for data-scarce learning of polymer properties
Ning Liu1, Siavash Jafarzadeh2, Brian Y Lattimer3
1Global Engineering and Materials Inc., Princeton, NJ, USA.
This study introduces a physics-based training pipeline to address data scarcity in material modeling with large language models (LLMs). The method uses synthetic data for pretraining, improving LLM accuracy for tasks like polymer flammability prediction.
Area of Science:
- Materials Science
- Artificial Intelligence
- Computational Modeling
Background:
- Large language models (LLMs) show potential for material modeling but require extensive data for accuracy.
- Acquiring sufficient experimental data for LLM fine-tuning is often limited and costly.
- Data scarcity poses a significant challenge for developing accurate LLM-based material models.
Purpose of the Study:
- To develop a physics-based training pipeline to overcome data scarcity in LLM material modeling.
- To enable accurate LLM training and fine-tuning even with limited experimental datasets.
- To improve the reliability and applicability of LLMs in material evaluation, analysis, and design.
Main Methods:
- A physics-based modeling framework generates abundant synthetic data for initial LLM alignment.
- A two-phase training strategy is employed: supervised pretraining with synthetic data, followed by fine-tuning with experimental data.
- The pipeline was validated using polymer flammability metrics, a domain with sparse experimental data.
Main Results:
- Supervised pretraining with synthetic data is crucial for achieving accurate fine-tuned LLMs.
- The physics-based pipeline effectively mitigates the negative impact of data scarcity.
- The approach demonstrated success in learning polymer flammability metrics.
Conclusions:
- Physics-based data generation is a viable strategy to address data scarcity in LLM material modeling.
- The proposed two-phase training pipeline enhances LLM accuracy and reliability.
- This methodology offers a pathway for more efficient and effective material design and analysis using LLMs.
Related Concept Videos
Polymers
Polymers: Molecular Weight Distribution
Polymer Classification: Crystallinity
Crystalline domains are the regions where polymer chains are aligned in an orderly manner and held together in proximity by intermolecular forces. For example, chains in the crystalline domains of polyethylene and nylon are bound together by van der Waals...
Polymer Classification: Architecture
Molecular Weight of Step-Growth Polymers
As the step-growth polymerization involves step-wise condensation of monomers, the molecular weight also builds up eventually. Consequently, high molecular weight polymers are obtained at the late stages of the polymerization, where 99% of monomers have been consumed.
The extent of the...
Step-Growth Polymerization: Overview
Many natural and synthetic polymers are produced by...

