Related Experiment Video
Updated: May 11, 2026

Electroantennographic Bioassay as a Screening Tool for Host Plant Volatiles
Published on: May 6, 2012
Accelerating VOC-OBP interaction screening via machine learning and molecular docking: Towards semiochemicals
Xaviera López-Cortés1, Nicolás Fernández2, Gabriel Lara2
1Department of Computer Sciences and Industries, Universidad Católica del Maule, Talca 3466706, Chile; Centro de Innovación en Ingeniería Aplicada (CIIA), Universidad Católica del Maule, Talca 3466706, Chile.
None:
This study presents a machine learning-based approach for predicting the binding affinity between volatile organic compounds (VOCs) and a category of odorant-binding proteins (OBPs) from moths to accelerate molecular screening and reduce experimental workload. A diverse set of regression models was evaluated, including ensemble methods (LightGBM, XGBoost, Gradient Boosting, Random Forest), kernel- based techniques (Support Vector Regressor), neural networks (Convolutional Neural Network), and a Bayesian linear model. The evaluation metrics included the coefficient of determination (R2), root mean square error (RMSE), and mean absolute error (MAE). Among all models, the LightGBM Regressor achieved the best results, with an R2 of 0.7101, RMSE of 0.2979, and MAE of 0.2200, outperforming all other approaches in predictive accuracy and error minimization. The results demonstrate that ensemble-based boosting algorithms are particularly well-suited for this task, effectively capturing complex, non-linear relationships in the data. In contrast, the Bayesian Ridge Regressor showed the weakest performance, highlighting the limitations of linear models in this context. Overall, the study illustrates the potential of artificial intelligence to transform molecular discovery pipelines, making them more efficient, cost-effective, and scalable. Furthermore, it lays the basis for future efforts related to the discovery of semiochemicals, key molecules for monitoring and control traps in integrated pest management. Finally, future work will focus on expanding training data and refining model architectures to enhance prediction quality and generalizability further.

