可解释机器学习和大数据挖掘,预测聚合物-MOF混合矩阵膜中的CO2分离
Hao Wan1,2, Yue Fang2, Min Hu2
1Guangzhou Key Laboratory for New Energy and Green Catalysis, School of Chemistry and Chemical Engineering, Guangzhou University, Guangzhou, 510006, P. R. China.
Advanced science (Weinheim, Baden-Wurttemberg, Germany)
|February 27, 2025
概括
高通量选和机器学习确定了最佳的混合矩阵膜 (MMM) 用于二氧化碳分离,超过性能限制. 预测模型有助于设计先进的MMM,以实现高效的气体分离.
科学领域:
- 材料科学 材料科学 材料科学
- 化学工程是化学工程的重要组成部分.
- 计算化学的计算化学
背景情况:
- 混合矩阵膜 (MMM) 对于气体分离至关重要,特别是对于二氧化碳 (CO2).
- 优化MMM性能需要了解聚合物和金属有机框架 (MOFs) 之间的复杂结构属性关系.
- 传统的实验查是耗时和昂贵的.
研究的目的:
- 通过计算选一个庞大的MMM数据库,以获得卓越的CO2分离性能.
- 使用机器学习开发准确的MMM性能预测模型.
- 在MMM中确定控制CO2分离的关键材料描述符.
主要方法:
- 在54,117个MMM (9个聚合物,6013个MOF) 的高通量虚拟选中.
- 机器学习,包括堆叠集团回归和沙普利增量解释 (SHAP),用于性能预测和特征重要性分析.
- 对CO2/CH4,CO2/N2,CO2/H2和CO2/O2二进制混合物的结构-性质关系的分析.
- 为MMM性能计算开发一个交互式桌面软件.
主要成果:
- 识别了超过罗伯森二氧化碳分离上限的MMM组合.
- 开发了一个高度准确的堆叠集团回归模型 (R2 = 0.96),具有出色的外推能力 (R2 = 0.95).
- 确定MOF孔腔尺寸和聚合物分数自由体积/密度是二氧化碳分离的关键特征.
结论:
- 计算选和机器学习为快速评估和设计高性能MMM提供了强大的方法.
- 转移学习对于预测二氧化碳分离性能和使用大型数据集设计新材料是有效的.
- 开发的软件为研究人员提供了高效的二氧化碳分离性能计算.
相关概念视频
Statistical Software for Data Analysis and Clinical Trials
478
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
478
Statistical Analysis System (SAS)
97
SAS, short for Statistical Analysis System, is a powerful data analysis, management, and visualization tool. Developed by the SAS Institute in the early 1970s, SAS has evolved into a comprehensive software suite used across various industries for statistical analysis, business intelligence, and predictive modeling.
Applications: SAS finds applications in numerous fields, including healthcare for clinical trial analysis, finance for risk assessment, marketing for customer data analysis, and...
Applications: SAS finds applications in numerous fields, including healthcare for clinical trial analysis, finance for risk assessment, marketing for customer data analysis, and...
97
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Introduction to R
224
R is a powerful software environment for statistical computing and graphics. Originating as an implementation of the S language, developed at Bell Laboratories, R has evolved into a robust, open-source statistical software favored by statisticians and data scientists worldwide. Its comprehensive suite includes data manipulation, calculation, and graphical display capabilities, making it versatile for data analysis and visualization. Its programming language is at the core of R's...
224
Statgraphics
99
Statgraphics is a comprehensive statistical software suite designed for both basic and advanced data analysis. Originating in 1980 at Princeton University under Dr. Neil W. Polhemus, it was one of the pioneering tools for statistical computing on personal computers, with its public release in 1982 marking an early milestone in data science software. Over the years, it has evolved into a robust platform for data science, offering tools for regression analysis, ANOVA, multivariate statistics,...
99
Outliers and Influential Points
3.9K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
3.9K


