Related Experiment Video
Updated: Sep 17, 2025

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Predictive Modeling of Yield Sooting Index Using Machine Learning with Uncertainty Estimation
Zied Hosni1, Xike Chen1, Sofiene Achour2,3
1University College London, Gower Street, London WC1E 6BT, United Kingdom.
Abstract:
This study explores the development of two predictive models for the yield sooting index (YSI) of various fuels using the advanced capabilities of machine learning (ML), particularly multilayer perceptron (MLP) networks. Quantitative structure-property relationship (QSPR) methodology, which connects molecular structures with fuel properties, enables accurate predictions of fuel behavior, including YSI, kinematic viscosity, ignition temperature, and cetane and octane numbers. By utilizing feature selection techniques such as Gini importance and genetic algorithms, we identified key molecular descriptors that significantly impact YSI. Remarkably, the genetic algorithm model outperformed the Gini importance model by effectively reducing autocorrelation among features, thereby enhancing the accuracy of predictions. The reliability of these models was further validated through uncertainty estimation at different significance levels, which provided deeper insights into their performance. Additionally, we identified a strong relationship between certain 2D matrix-based descriptors and YSI, which offers a fresh perspective on predicting fuel properties. This comprehensive approach, incorporating rigorous data preprocessing, feature selection, and hyperparameter tuning, demonstrated the robustness of the developed predictive models. This work highlights the potent synergy between ML and QSPR theory in advancing the prediction of fuel properties. It not only refines current predictive models but also sets the stage for future computational advancements in fuel research, which contributes to the broader goal of developing sustainable and efficient fuel alternatives.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Propagation of Uncertainty from Systematic Error
Uncertainty: Confidence Intervals
Propagation of Uncertainty from Random Error
Survival Tree
Building a Survival Tree
Constructing a...

