Related Experiment Video
Updated: Aug 16, 2025

Experimental Protocol for Manipulating Plant-induced Soil Heterogeneity
Published on: March 13, 2014
Compositionality, sparsity, spurious heterogeneity, and other data-driven challenges for machine learning algorithms
Sebastiano Busato1, Max Gordon1, Meenal Chaudhari1
1Department of Electrical and Computer Engineering, North Carolina State University, Raleigh, USA; NC Plant Sciences Initiative, North Carolina State University, Raleigh, USA.
Abstract:
The plant-associated microbiome is a key component of plant systems, contributing to their health, growth, and productivity. The application of machine learning (ML) in this field promises to help untangle the relationships involved. However, measurements of microbial communities by high-throughput sequencing pose challenges for ML. Noise from low sample sizes, soil heterogeneity, and technical factors can impact the performance of ML. Additionally, the compositional and sparse nature of these datasets can impact the predictive accuracy of ML. We review recent literature from plant studies to illustrate that these properties often go unmentioned. We expand our analysis to other fields to quantify the degree to which mitigation approaches improve the performance of ML and describe the mathematical basis for this. With the advent of accessible analytical packages for microbiome data including learning models, researchers must be familiar with the nature of their datasets.
More Related Videos
Related Concept Videos
Modern Molecular Taxonomy
Mechanistic Models: Compartment Models in Individual and Population Analysis
Applications of Molecular Taxonomy
The Roles of Bacteria and Fungi in Plant Nutrition
Introduction to Plant Diversity
Environmental Applications of Microorganisms

