Related Experiment Video
Updated: Jul 14, 2026

Quasistatic Mechanical Testing for Computer-Aided Design and Manufacturing Occlusal Veneers Cemented to Milled Dentin Analog Material
Published on: December 20, 2024
Clever materials: when models identify good materials for the wrong reasons
Kevin Maik Jablonka1,2,3,4
1Laboratory of Organic and Macromolecular Chemistry (IOMC), Friedrich Schiller University Jena, Humboldtstrasse 10, 07743 Jena, Germany. mail@kjablonka.com.
Abstract:
Machine learning can accelerate materials discovery. Models perform impressively on many benchmarks. However, strong benchmark performance does not imply that a model has learned chemistry. I test a concrete alternative hypothesis: that property prediction can be driven by bibliographic confounding. Across five tasks spanning MOFs (thermal and solvent stability), perovskite solar cells (efficiency), batteries (capacity), and TADF emitters (emission wavelength), models trained on standard chemical descriptors predict author, journal, and publication year well above chance. When these predicted metadata ("bibliographic fingerprints") are used as the sole input to a second model, performance is sometimes competitive with conventional descriptor-based predictors. These results show that many datasets do not rule out non-chemical explanations of success. Progress requires routine falsification tests (e.g., group/time splits and metadata ablations), datasets designed to resist spurious correlations, and explicit separation of two goals: predictive utility versus evidence of chemical understanding.
Related Concept Videos
Bending of Material: Problem Solving
Molecular Models
Polymer Classification: Architecture
Stereotype Content Model
Polymer Classification: Stereospecificity
Predicting Products: Substitution vs. Elimination
The following factors can influence the mechanisms competing against each other:

