Related Experiment Video
Updated: Jul 10, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Towards a framework for interoperability and reproducibility of predictive models
Al Rahrooh1, Anders O Garlid1, Kelly Bartlett1
1Medical & Imaging Informatics (MII) Group, University of California Los Angeles (UCLA), Los Angeles, CA, USA.
This study introduces an Automated Metadata Pipeline (AMP) to standardize machine learning (ML) model reproducibility in healthcare. AMP uses extended Predictive Model Markup Language (PMML) to ensure models are shareable and comparable.
Area of Science:
- Biomedical informatics
- Machine learning applications in healthcare
- Computational biology
Background:
- Standardized methodologies for developing and deploying machine learning (ML) models in biomedical research and healthcare are currently lacking.
- Existing tools for model replication do not provide a unifying blueprint, hindering scientific reproducibility due to unclear assumptions, preprocessing steps, and test metrics.
- Challenges in generalizability and transportability of ML models remain significant issues in the field.
Purpose of the Study:
- To facilitate scientific reproducibility in biomedical machine learning.
- To present the Automated Metadata Pipeline (AMP) as a key component of the PREdictive Model Index and Exchange REpository (PREMIERE) platform.
- To enable the conversion of predictive ML models into an extended PMML format for enhanced interoperability and reproducibility.
Main Methods:
- Development of the Automated Metadata Pipeline (AMP) building upon the Predictive Model Markup Language (PMML).
- Conversion of predictive ML models into extended PMML files.
- Autocompletion of an ML-based checklist to assess model elements for interoperability and reproducibility.
- Demonstration of the pipeline on multiple test cases using three different ML algorithms and health-related datasets.
Main Results:
- Successful conversion of ML models into extended PMML files using the AMP.
- Demonstrated assessment of model elements for interoperability and reproducibility via an autocompleted checklist.
- Validation of the pipeline across diverse ML algorithms and health datasets.
Conclusions:
- The Automated Metadata Pipeline (AMP) provides a foundational solution for enhancing the reproducibility of predictive ML models in healthcare.
- The extended PMML format facilitates better model sharing, comparison, and understanding of generalizability and transportability.
- This work paves the way for more robust and reliable ML applications in biomedical research and clinical practice.
More Related Videos
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Improving Translational Accuracy
Statistical Software for Data Analysis and Clinical Trials
Correlation and Regression
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...

