Related Experiment Video
Updated: Jul 10, 2026

O-cresol Concentration Online Measurement Based On Near Infrared Spectroscopy Via Partial Least Square Regression
Published on: November 8, 2019
Functional clustering as a correction framework for regression models under small-data constraints: predicting
1FSBIS Institute of Physiologically Active Compounds of the Russian Academy of Sciences, Russian Academy of Sciences, 1, Severny proezd, Chernogolovka 142432, Moscow Region, Russian Federation. tolbin@ipac.ac.ru.
Functional clustering improves inaccurate regression models for materials science, reducing prediction errors from over 30% to under 25%. This method enhances accuracy for small datasets, aiding in property prediction.
Area of Science:
- Materials Chemistry
- Computational Chemistry
- Data Science
Background:
- Regression models often perform poorly with limited data in materials science.
- Accurate prediction of material properties is crucial for accelerating discovery.
- Existing methods struggle with small sample sizes and high initial error rates.
Purpose of the Study:
- To develop a method for correcting poorly performing regression models under small-data constraints.
- To improve the accuracy of predicting optical limiting properties in phthalocyanines.
- To establish a quantitative applicability domain and interpretable rules for new compound prediction.
Main Methods:
- Functional clustering that jointly optimizes cluster assignments and local correction functions.
- Bayesian information criterion for determining the optimal number of clusters.
- Leave-one-out cross-validation for assessing generalizability within clusters.
- Decision trees for generating interpretable IF-THEN rules for new compound classification.
Main Results:
- Reduced global prediction error from 32-140% to 10-25% mean absolute percentage error.
- Demonstrated generalizability within well-populated clusters (≥5 compounds) with median errors of 14-36%.
- Identified cluster instability for small clusters (3-4 compounds), defining a quantitative applicability domain.
- Feature importance analysis revealed that cluster assignment relies on descriptors distinct from raw regression inputs, uncovering latent physicochemical structure.
Conclusions:
- Functional clustering is a generalizable and effective method for correcting inaccurate regression models in small-data scenarios within materials chemistry.
- The method provides accurate property predictions and uncovers underlying structure-property relationships.
- Open-source code enables researchers to apply this technique for prescreening compound libraries, accelerating materials discovery.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Variables Affecting Phosphorescence and Fluorescence
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Model Approaches for Pharmacokinetic Data: Compartment Models
Two primary types of compartment models are recognized: mammillary and catenary. The more...

