Related Experiment Video
Updated: Aug 6, 2025

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
A real data-driven simulation strategy to select an imputation method for mixed-type trait data
Jacqueline A May1, Zeny Feng2, Sarah J Adamowicz1
1Department of Integrative Biology & Biodiversity Institute of Ontario, University of Guelph, Guelph, Ontario, Canada.
Selecting the best method to fill in missing trait data is crucial for biological analyses. A data-driven simulation using squamate traits found random forests with phylogenetic information to be most effective for imputation.
Area of Science:
- Evolutionary Biology
- Bioinformatics
- Comparative Genomics
Background:
- Missing observations in biological trait datasets hinder analyses across various disciplines.
- Existing imputation methods yield mixed results, necessitating a framework for selecting appropriate techniques for diverse, real-world datasets.
- Trait datasets often contain mixed data types (categorical, count, continuous), complicating imputation strategies.
Purpose of the Study:
- To develop and validate a real data-driven simulation strategy for selecting the optimal imputation method for mixed-type trait datasets.
- To evaluate the performance of candidate imputation methods, including mean/mode, k-nearest neighbour, random forests, and MICE, with and without phylogenetic information.
Main Methods:
- A squamate trait dataset was used as a target, with missing data simulated under missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR) mechanisms.
- Imputation was performed using candidate methods, incorporating phylogenetic information from nuclear, mitochondrial, or multigene trees.
- Performance was assessed using mean squared error for numerical traits and proportion falsely classified rates for categorical traits.
Main Results:
- The random forest method, enhanced with a nuclear-derived phylogeny, demonstrated the lowest error rates across most traits.
- Imputed datasets more accurately reflected the original data's characteristics and distributions compared to complete-case datasets.
- Phylogenetic information did not consistently improve performance for all traits or scenarios, highlighting the need for careful method selection.
Conclusions:
- A real data-driven simulation strategy is effective for selecting suitable imputation methods for mixed-type trait datasets.
- Random forests combined with appropriate phylogenetic data offer a robust approach for trait data imputation in evolutionary biology.
- Caution is advised, as the utility of phylogenetic information in imputation varies by trait and missingness mechanism.
Related Concept Videos
Multiple Allele Traits
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Survival Tree
Building a Survival Tree
Constructing a...
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...

