Related Experiment Video
Updated: Aug 8, 2026

Characterization of Complex Systems Using the Design of Experiments Approach: Transient Protein Expression in Tobacco as a Case Study
Published on: January 31, 2014
Construction and optimization of a genomic selection model for total sugar content in tobacco
Abstract:
Total sugar is one of the chemical indicators related to tobacco quality. However, due to its composition of various oligosaccharides, disaccharides, and polysaccharides, total sugar lacks a clear genetic target, making conventional breeding approaches for its improvement particularly challenging. Therefore, conducting genomic selection (GS) studies on total sugar content holds significant importance for the development of high-sugar tobacco varieties. In this study, 2,604 tobacco germplasm accessions with broad genetic diversity, sourced from the National tobacco Germplasm Repository, were used as experimental materials. High-throughput sequencing technologies were employed to perform comprehensive genetic evaluation and construct genomic selection models for total sugar content. Sixteen mainstream genomic prediction models were assessed through five-fold cross-validation to develop a high-accuracy prediction framework. Among these models, the Gradient Boosting Machine (GBM) achieved the highest prediction accuracy for total sugar content (0.85), followed by the rrBLUP model (0.81). Comparative analysis of computational speed and resource consumption revealed that GBM maintained rapid computation and low resource usage even in large sample sizes, demonstrating strong stability and superior performance. Considering all factors, GBM was preliminarily identified as the optimal model for predicting total sugar content in tobacco. The application of high-accuracy genomic prediction is expected to overcome the challenges of phenotypic evaluation in breeding programs and significantly enhance the efficiency of selecting high-sugar tobacco varieties.
