Related Experiment Video
Updated: Jul 25, 2025

06:19
Integration of Animal Behavioral Assessment and Convolutional Neural Network to Study Wasabi-Alcohol Taste-Smell Interaction
Published on: August 16, 2024
456
Developing high-dimensional machine learning models to improve generalization ability and overcome data insufficiency
Xiao-Yan Huang1, Tian-Jie Ao1, Xue Zhang1
1State Key Laboratory of Microbial Metabolism, Joint International Research Laboratory of Metabolic & Developmental Sciences, School of Life Sciences and Biotechnology, Shanghai Jiao Tong University, Shanghai 200240, China.
Bioresource Technology
|June 23, 2023
Summary
Accurate machine learning models boost biorefineries. This study enhances model generalization and data efficiency for mixed sugar fermentation simulation, improving prediction accuracy and reducing data needs.
Area of Science:
- Biotechnology
- Machine Learning
- Biorefinery Processes
Background:
- Accurate machine learning models are crucial for advancing biorefinery efficiency.
- Simulating mixed sugar fermentation presents challenges due to data limitations and model generalization.
Purpose of the Study:
- To develop a strategy for enhancing machine learning model generalization and overcoming data insufficiency in mixed sugar fermentation simulation.
- To improve the accuracy and applicability of artificial neural network (ANN) models in biorefinery contexts.
Main Methods:
- Utilized multiple inputs single output (MISO) and multiple inputs multiple outputs (MIMO) models for fermentation simulation.
- Developed a consensus yeast (CY) model by consolidating data from multiple yeast strains.
- Adjusted pretrained CY models to assess performance with reduced datasets.
Main Results:
- MISO models with combined initial glucose, xylose, and time inputs showed superior generalization compared to time-only inputs.
- MIMO models achieved high accuracy with an average R² of 0.99.
- The CY model achieved an R² of 0.90, and adjusted models retained high accuracy (R² 0.95 for yeast, 0.93 for bacterial) with over 50% less data.
Conclusions:
- The proposed strategy significantly enhances model generalization and addresses data scarcity in fermentation simulations.
- This approach expands the applicability of ANN models in biorefineries and reduces data curation costs.

