Related Experiment Video
Updated: Jan 13, 2026

Author Spotlight: UAV Remote Sensing for Efficient Invasive Plant Biomass Estimation
Published on: February 9, 2024
A novel clustering-regression machine learning framework for biomass classification and biochemical composition
Jiaxin Gao1, Weijin Zhang1, Lijian Leng1
1School of Energy Science and Engineering, Central South University, Changsha 410083, China; Xiangjiang Laboratory, Changsha 410205, China.
Abstract:
Biomass elemental and biochemical compositions determine its conversion behavior and utilization potential. However, a standardized classification system based on these intrinsic characteristics is lacking, and different biomass types follow distinct processing pathways. In comparison, biomass biochemical analysis is time-consuming and costly, limiting its large-scale evaluation and application. To address these challenges, this study developed a novel clustering-regression machine learning (ML) framework that predicts the main biochemical components of different biomass types based on their elemental composition. First, Principal Component Analysis (PCA) achieved effective dimensionality reduction, with a total explained variance of 93.2 %. Using a PCA-assisted clustering model, biomass was effectively clustered into lipid-/protein-rich and lignocellulosic groups with a high silhouette score of 0.605. Subsequently, regression model 1 predicted protein and lipid contents in lipid-/protein-rich biomass, achieving average test R2 of 0.88 and average validation R2 of 0.81. For lignocellulosic biomass, regression model 2 predicted fibres and lignin contents, achieving average test R2 of 0.77 and average validation R2 of 0.66. Clustering and regression models all underwent thorough validation and testing, showing strong generalization. To make the integrated clustering-regression framework more user-friendly, a software application is established. Users can input elemental composition to identify biomass clusters and predict the corresponding main component contents. This study provides a reliable predictive tool for screening suitable feedstocks for targeted product manufacturing.
More Related Videos
11:14Rapid High-throughput Species Identification of Botanical Material Using Direct Analysis in Real Time High Resolution Mass Spectrometry
Published on: October 2, 2016
11:31High-throughput Screening of Recalcitrance Variations in Lignocellulosic Biomass: Total Lignin, Lignin Monomers, and Enzymatic Sugar Release
Published on: September 15, 2015
Related Concept Videos
Classification of Elements and Compounds
Compounds are pure substances composed of two or more elements in fixed, definite proportions. Compounds are classified as ionic or molecular (covalent) based on the bonds...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Classification of Systems-II
Classification of Titrimetric Analysis Based on Reaction Types
Titrations between an acid and a base lead to neutralization reactions that form...