Related Experiment Video
Updated: Feb 28, 2026

Field Identification of Matricaria chamomilla using a Portable qPCR System
Published on: October 10, 2020
Origin Identification of Scutellariae radix Based on Multidimensional Quality Indicators and Machine Learning
Xiao-Lu Liu1, Tong Zhu1,2, Ming-Yue Zhang2
1School of Chinese Materia Medica, Chongqing University of Chinese Medicine, Bishan, Chongqing 402760, China.
None:
This study aims to establish an origin identification method for Scutellariae radix that integrates multidimensional quality indicators and machine learning algorithms, enabling accurate and rapid traceability of Scutellariae radix medicinal materials from four production areas: Hebei (HB), Shanxi (SX), Shaanxi (SAX), and Chengde (CD). The study collected a total of 43 batches of Scutellariae radix samples from the aforementioned origins. It systematically measured 12 key quality indicators covering flavonoids, physicochemical parameters, chromaticity values, and biological activity. These specifically include four flavonoid components: baicalin, wogonoside, baicalein, and wogonin; three physicochemical parameters: moisture content, ash content, and alcohol-soluble extract; four chromaticity values: L*, a*, b*, and ΔE; and in vitro anti-inflammatory activity (IC50 value for NO clearance). On the basis of these parameters, in this study there were five machine learning models constructed based on the following algorithms and methods: Random Forest (RF), Extreme Learning Machine (ELM), Backpropagation Neural Network (BP), and Radial Basis Function Neural Network (RBF). A comparative analysis was conducted to evaluate the origin identification performance of each model. The results indicate significant differences (p < 0.05) in the contents of baicalin, wogonoside, L*, a*, b*, ΔE, and alcohol-soluble extract among Scutellariae radix from different origins. The comparative analysis of four machine learning models reveals that RF outperforms ELM, BP, and RBF in multiclass classification, achieving a test accuracy of 75% and consistent precision, recall, and F1-score of 79.17%. In contrast, the three neural networks attain only 66.67% test accuracy, with RBF showing high precision but low recall, ELM delivering moderate performance, and BP performing poorly. These results underscore the strength of ensemble methods like RF in small-sample settings, where they mitigate overfitting and enhance generalization, whereas neural networks struggle with limited data. We therefore recommend RF for deployment under current data constraints and suggest future work should focus on data expansion, especially for under-performing classes, along with hyperparameter tuning to further improve classification.

