Related Experiment Videos
When high accuracy misleads in literature-derived machine learning for deep eutectic solvent recommendation
Hakim Faraji1, Juan A Mendez-Perez2, Ruth Rodríguez-Ramos1,3
1Departamento de Química, Área de Química Analítica, Facultad de Ciencias, Universidad de La Laguna (ULL) Avenida Astrofísico Francisco Sánchez s/n 38206 San Cristóbal de La Laguna Tenerife Spain hafaraji@ull.edu.es.
RSC Advances
|August 5, 2026
Summary
Machine learning (ML) shows promise for selecting deep eutectic solvents (DESs) but struggles with literature data. Improved accuracy in DES extraction models doesn't guarantee reliable recommendations without better data structure.
Area of Science:
- Analytical Chemistry
- Computational Chemistry
Background:
- Deep eutectic solvents (DESs) are versatile for analytical sample preparation.
- Selecting optimal DES systems is challenging due to vast design space and empirical methods.
- Machine learning (ML) offers a data-driven approach, but its efficacy with literature data is uncertain.
Purpose of the Study:
- To evaluate the reliability of ML models for predicting DES performance using literature-derived data.
- To assess if improved predictive accuracy translates to practical recommendation capabilities.
- To identify data quality and structure limitations impacting ML model performance in DES selection.
Main Methods:
- Reconstructed a literature-derived dataset of 757 DES-based pesticide extraction experiments from 94 studies.
- Implemented a leakage-controlled, DOI-grouped validation framework.
- Developed and compared a hybrid ML model (classification and preference learning) against a baseline.
Main Results:
- The hybrid ML model improved record-level classification but did not yield stable group-level ranking benefits.
- Performance degraded in restricted comparable groups, with the hybrid model often underperforming.
- Baseline models captured significant structure, but the hybrid model offered no consistent additional advantage.
Conclusions:
- Predictive accuracy gains in ML models for DES extraction do not automatically ensure reliable chemical recommendations from literature data.
- Sparse comparative data, study-specific targets, and limited descriptors hinder ML model utility.
- Future ML applications require data-centric evaluation, standardized reporting of comparable candidates, and structured metadata.