Related Experiment Video
Updated: Aug 22, 2026

Clean Sampling and Analysis of River and Estuarine Waters for Trace Metal Studies
Published on: July 1, 2016
Objective-Dependent Active Learning and Calibrated Uncertainty for Sample-Efficient Discovery of Rare-Earth
1Lawrence E. Elkins High School, Missouri City, Texas 77459, United States.
None:
Measuring distribution coefficients (log D) for rare-earth solvent extraction is slow and expensive, making sample-efficient machine learning attractive. However, the practical questions of which acquisition strategy to use, whether the attendant uncertainty estimates can be trusted, and how well a model actually generalizes to new chemistry are rarely answered together on real data. We benchmark active learning (AL) on 1,202 experimental log D measurements for lanthanide extraction (93 ligands, 14 lanthanides, 45 literature sources), with prospective validation on four independently synthesized ligands. Five findings emerge, each based on real data with multiseed statistics. (i) A plain random forest reaches R 2 = 0.94/RMSE = 0.33 on the published row-level validation split, outperforming the previously reported deep network (R 2 = 0.85). However, under leakage-controlled splits that forbid a ligand, source, or structural family from appearing in both train and test, the estimate collapses (grouped-CV R 2 = 0.20 to -0.03), so the in-distribution number is strongly optimistic, and realistic deployability is far lower. (ii) The optimal AL acquisition is objective-dependent: exploitation recovers all top row-level extractants (top-10 recall = 1.0) but yields a biased, poorly predictive model, whereas exploration/random gives the best predictive model but finds few of the best. No strategy dominates both, and this survives a chemically realistic ligand-batch acquisition unit and reproduces on an independent 4,200-compound log D benchmark. (iii) The discovery ranking of strategies inverts with the definition of a "top" system: exploitation dominates for condition-specific rows, but broad sampling already recovers most top ligands. (iv) For prediction, AL provides no advantage over random selection. (v) Conformal calibration is essential for trustworthy uncertainty (an 8-fold in-distribution error reduction) but does not improve AL acquisition, which depends only on the uncertainty ranking. Under distribution shift, global conformal leaves subgroup imbalance that group-conditional (Mondrian) calibration reduces, albeit at the cost of reordering the ranking. We compare five model/uncertainty estimators, add early recognition discovery metrics, and distill the results into a practical decision guide, releasing all code and data.
More Related Videos
10:22Split Point Analysis and Uncertainty Quantification of Thermal-Optical Organic/Elemental Carbon Measurements
Published on: September 7, 2019
10:31Detection and Recovery of Palladium, Gold and Cobalt Metals from the Urban Mine Using Novel Sensors/Adsorbents Designated with Nanoscale Wagon-wheel-shaped Pores
Published on: December 6, 2015
Related Concept Videos
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an organic...
Extraction: Advanced Methods
Calculating Equilibrium Concentrations
A more...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Atomic Absorption Spectroscopy: Lab
Solutions containing organic solvents, such as low-molecular-mass alcohols, esters, or ketones, enhance absorbances by increasing nebulizer...
The Equilibrium Binding Constant and Binding Strength