Interpretable Machine-Learning and Big Data Mining to Predict the CO2 Separation in Polymer-MOF Mixed Matrix
Hao Wan1,2, Yue Fang2, Min Hu2
1Guangzhou Key Laboratory for New Energy and Green Catalysis, School of Chemistry and Chemical Engineering, Guangzhou University, Guangzhou, 510006, P. R. China.
Abstract:
Mixed matrix membranes (MMMs) are renowned for their exceptional gas separation capabilities. In this work, high-throughput computing screening and machine learning are employed to evaluate the CO2 separation performance of 54117 MMMs composed of 9 polymers and 6013 metal-organic frameworks (MOFs). The structure-property relationships of MMMs are analyzed for 4 binary mixtures (CO2/X, X = CH4, N2, H2, O2), and the best-performing combinations of MOFs and polymers are found, with which the MMM performance exceeded the Robeson's upper limit. Then, a stacked ensemble regression model with high accuracy (average R2 = 0.96) is trained, demonstrating excellent extrapolation capability (R2 = 0.95) for new MMMs containing 6FDA-DAM. Furthermore, by utilizing Shapley Additive Explanations and data segmentation, it is identified that the pore limit diameter and largest cavity diameter in MOF features and the fractional free volume and density in polymer features are of paramount importance. Two extrapolation methods are compared and found that transfer learning is better for predicting CO2 separation performance in MMMs and designing new materials with large datasets. Finally, an interactive desktop software is developed to assist researchers in rapidly and accurately calculating the CO2 separation performance of MMMs. This work presents a novel approach for the rapid evaluation of high-quality MMMs and the efficient calculation of gas permeation rates within membranes.
More Related Videos
Related Concept Videos
Statistical Software for Data Analysis and Clinical Trials
Statistical Analysis System (SAS)
Applications: SAS finds applications in numerous fields, including healthcare for clinical trial analysis, finance for risk assessment, marketing for customer data analysis, and...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Introduction to R
Statgraphics
Outliers and Influential Points


