Related Experiment Video
Updated: Jan 31, 2026

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
Published on: November 22, 2019
A heuristic approach to handling missing data in biologics manufacturing databases
Jeanet Mante1, Nishanthi Gangadharan2, David J Sewell2
1Pembroke College, Cambridge, UK.
This study evaluated data imputation methods for historical biomanufacturing datasets. Mean imputation suits simple data, while regression imputation is effective for complex data with under 30% missing values, depending on data characteristics.
Area of Science:
- Biotechnology
- Data Science
- Process Engineering
Background:
- Biologics manufacturing generates vast historical datasets.
- Inter-batch variability and missing data are common due to diverse experiments and technical issues.
- Data pre-processing is crucial for accurate analysis and predictive modeling.
Purpose of the Study:
- To investigate the efficiency of mean imputation and multivariate regression for handling missing data in biomanufacturing datasets.
- To evaluate the performance of these imputation methods using symbolic regression and Bayesian non-parametric models.
- To identify key factors influencing the selection of appropriate imputation techniques.
Main Methods:
- Comparison of mean imputation and multivariate regression imputation.
- Application of symbolic regression models for data processing.
- Utilization of Bayesian non-parametric models for evaluating imputation performance.
- Analysis of missing data mechanisms (Missing Completely At Random, Missing At Random, Missing Not At Random).
Main Results:
- Mean imputation is efficient for smooth, non-dynamical datasets.
- Regression imputation effectively preserves data distribution and standard deviation for dynamical datasets with <30% missing data.
- The nature of missing data (MCAR, MAR, MNAR) is critical for method selection.
Conclusions:
- The choice of imputation method depends on dataset characteristics and the nature of missing data.
- Appropriate pre-processing enhances the reliability of data mining in biomanufacturing.
- Understanding missing data mechanisms is key to effective data imputation and predictive modeling.
Related Concept Videos
The Availability Heuristic
The Representativeness Heuristic
The Anchoring-and-Adjustment Heuristic
Heuristics
People often rely on heuristics when faced with an overload of information, limited time, low importance of the decision, limited information, or when a heuristic readily comes to mind. For...
Model Approaches for Pharmacokinetic Data: Compartment Models
Two primary types of compartment models are recognized: mammillary and catenary. The more...
Model Approaches for Pharmacokinetic Data: Physiological Models

