Prediction of sweetness by multilinear regression analysis and support vector machine
Min Zhong1, Yang Chong, Xianglei Nie
1State Key Laboratory of Chemical Resource Engineering, Dept. of Pharmaceutical Engineering, Beijing Univ. of Chemica Technology, Beijing, China.
Journal of Food Science
|August 7, 2013
Summary
This study developed two quantitative models to predict compound sweetness, aiding the food additive industry. Machine learning approaches accurately forecast sweetness based on molecular structure, supporting established taste theories.
Area of Science:
- Computational Chemistry
- Food Science
- Cheminformatics
Background:
- Sweetness is a key characteristic for food additives.
- Predicting sweetness quantitatively is crucial for industry applications.
- Understanding structure-sweetness relationships can guide new sweetener development.
Purpose of the Study:
- To develop and validate quantitative models for predicting the logarithm of sweetness (logSw).
- To identify molecular descriptors that influence sweetness.
- To assess the performance of multilinear regression (MLR) and support vector machine (SVM) for sweetness prediction.
Main Methods:
- A dataset of 320 compounds with known sweetness values was compiled.
- Compounds were characterized by 12 molecular descriptors.
- The dataset was split into training (214 compounds) and testing (106 compounds) sets.
- MLR and SVM models were employed to predict logSw.
Main Results:
- Both MLR and SVM models achieved high predictive accuracy on the test set.
- Correlation coefficients of 0.87 (MLR) and 0.88 (SVM) were obtained.
- Identified molecular descriptors offer structural insights into sweetness perception.
Conclusions:
- Quantitative structure-activity relationship (QSAR) models can effectively predict compound sweetness.
- The findings support the established AH/B System model of taste perception.
- These models can aid in the design and selection of novel sweetening agents.
Related Concept Videos
Multiple Regression
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Regression Toward the Mean
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when researchers try to extrapolate results...
Multi-input and Multi-variable systems
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
Classification of Signals
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Correlation and Regression
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a negative...
