Related Experiment Video
Updated: May 18, 2026

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Introducing uncertainty in predictive modeling--friend or foe?
1AstraZeneca R&D Södertälje, Sweden. ulf.norinder@astrazeneca.com
Journal of Chemical Information and Modeling
|October 9, 2012
Summary
Introducing uncertainty into chemical descriptors does not significantly impair predictive performance for decision tree ensembles. Elaborate methods for handling descriptor uncertainty are unnecessary with this modeling approach.
Area of Science:
- Computational chemistry
- Cheminformatics
- Machine learning in chemistry
Background:
- Chemical descriptors are crucial for predictive modeling in cheminformatics.
- Uncertainty in descriptor values can arise from experimental or computational limitations.
- The impact of descriptor uncertainty on predictive model performance requires thorough investigation.
Purpose of the Study:
- To assess the effect of varying degrees and types of uncertainty in chemical descriptors on the predictive performance of decision tree ensembles.
- To evaluate different strategies for handling uncertainty within decision tree ensemble models.
- To determine if complex uncertainty handling methods offer advantages over simpler approaches.
Main Methods:
- Systematic introduction of uncertainty into chemical descriptors across 16 public datasets.
- Application of state-of-the-art decision tree ensemble methods for predictive modeling.
- Evaluation of various strategies for managing uncertainty in descriptor data.
- Comparison of model performance using single-point samples versus full uncertainty distributions.
Main Results:
- Predictive performance of decision tree ensembles remained largely unimpaired despite significant uncertainty introduced into chemical descriptors.
- Practical utility of predictive models was not substantially compromised by descriptor uncertainty.
- Ensemble models performed comparably whether trained on single-point samples or full distributions of uncertain values.
Conclusions:
- Decision tree ensembles exhibit robustness towards uncertainty in chemical descriptors.
- There is no practical benefit to employing complex uncertainty handling techniques for chemical descriptors when using decision tree ensembles.
- The findings suggest that simpler data handling strategies are sufficient for maintaining predictive accuracy in the presence of descriptor uncertainty.
Related Concept Videos
Uncertainty: Overview
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
Propagation of Uncertainty from Random Error
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
Uncertainty: Confidence Intervals
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor 't,' or...
Prediction Intervals
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
Propagation of Uncertainty from Systematic Error
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this particular...
Uncertainty in Measurement: Accuracy and Precision
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.