Related Experiment Video
Updated: Aug 17, 2025

09:32
Cooling Rate Dependent Ellipsometry Measurements to Determine the Dynamics of Thin Glassy Films
Published on: January 26, 2016
8.2K
Machine Learning with Enormous "Synthetic" Data Sets: Predicting Glass Transition Temperature of Polyimides Using
Igor V Volgin1, Pavel A Batyr2, Andrey V Matseevich3
1Institute of Macromolecular Compounds of the Russian Academy of Sciences (IMC RAS), St. Petersburg 199004, Russian Federation.
ACS Omega
|December 12, 2022
Summary
Machine learning predicts polymer thermal properties using graph convolutional neural networks (GCNNs). Transfer learning with synthetic data significantly improves accuracy for predicting polyimide glass transition temperatures.
Area of Science:
- Materials Science
- Computational Chemistry
- Machine Learning
Background:
- Predicting polymer thermal properties, like glass transition temperature (Tg), is crucial for material design.
- Traditional methods like quantitative structure-property relationship (QSPR) have limitations in accuracy and scope.
- Machine learning (ML) offers a powerful approach to establish complex structure-property relationships in polymers.
Purpose of the Study:
- To develop and validate a graph convolutional neural network (GCNN) for predicting the glass transition temperature (Tg) of polyimides (PIs).
- To investigate the efficacy of transfer learning using large synthetic datasets for pretraining GCNN models.
- To compare the performance of GCNNs trained on polymer-specific data versus general chemical data.
Main Methods:
- Development of a GCNN model for predicting PI Tg.
- Creation of a large synthetic dataset (>6 million PI repeating units) with theoretically calculated Tg values using Askadskii's QSPR.
- Implementation of a transfer learning strategy, pretraining on synthetic data and fine-tuning on a smaller experimental dataset (214 PIs).
- Comparison with GCNNs pretrained on the QM9 small molecule database.
Main Results:
- The GCNN model trained with transfer learning achieved a mean absolute error (MAE) of ~20 K for PI Tg prediction, outperforming Askadskii's QSPR (33 K).
- Pretraining on the synthetic polymer dataset resulted in significantly better performance (MAE ~20 K) compared to pretraining on the QM9 dataset (MAE ~41 K).
- The study highlights the importance of using polymer-specific data for pretraining deep learning models to mitigate the 'reality gap'.
Conclusions:
- The developed GCNN methodology, leveraging transfer learning and large synthetic polymer datasets, provides accurate predictions of polyimide glass transition temperatures.
- Using polymer-specific data for pretraining is superior to using general chemical data for developing effective ML models for polymer property prediction.
- The approach is versatile and can be extended to predict other properties of various polymers and copolymers.

