Related Experiment Video
Updated: Oct 31, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Ensuring the Robustness and Reliability of Data-Driven Knowledge Discovery Models in Production and Manufacturing.
Shailesh Tripathi1, David Muhr1, Manuel Brunner1
1Production and Operations Management, University of Applied Sciences Upper Austria, Linz, Austria.
The Cross-Industry Standard Process for Data Mining (CRISP-DM) framework needs extensions for robust data science. A new Generalized Cross-Industry Standard Process for Data Science (GCRISP-DS) framework enhances model reusability and addresses data issues.
Area of Science:
- Data Science
- Machine Learning
- Knowledge Discovery
Background:
- The Cross-Industry Standard Process for Data Mining (CRISP-DM) is a key framework for data analytics and machine learning in production and manufacturing.
- Practical implementation of CRISP-DM faces challenges with data and model development, requiring industry-specific solutions.
Purpose of the Study:
- To review the CRISP-DM framework and its extensions.
- To introduce a novel Generalized Cross-Industry Standard Process for Data Science (GCRISP-DS) framework.
- To enhance the flexibility and robustness of data science processes.
Main Methods:
- Detailed review of the CRISP-DM model.
- Summarization of existing CRISP-DM extensions.
- Development of the GCRISP-DS framework with dynamic phase interactions.
Main Results:
- The GCRISP-DS framework allows dynamic interactions between phases to address data and model issues.
- It emphasizes business understanding and data quality for achieving business objectives.
- The framework enhances model improvements and reusability while minimizing robustness issues.
Conclusions:
- The GCRISP-DS framework offers a customizable and flexible approach to data science.
- It effectively addresses practical challenges in implementing data-driven models.
- This enhanced framework supports the achievement of higher business objectives through robust data science practices.
More Related Videos
07:35A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
20:24Characterization of Complex Systems Using the Design of Experiments Approach: Transient Protein Expression in Tobacco as a Case Study
Published on: January 31, 2014
Related Concept Videos
Distribution Reliability and Automation
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Production Efficiency
Reliability and Validity