Related Experiment Video
Updated: Jul 2, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Bridging interpretable machine learning and large language models through direct representative selection and
Tomomi Shimazaki1, Masanori Tachikawa1
1Quantum Chemistry Division, Yokohama City University, Seto 22-2, Kanazawa-Ku, Yokohama 236-0027, Kanagawa, Japan. tshima@yokohama-cu.ac.jp.
Abstract:
In this study, we combined an interpretable machine learning (ML) framework with a large language model (LLM) to investigate structure-reactivity trends in an acrylate/methacrylate radical reaction dataset constructed from density functional theory calculations. For the ML component, we employed modified convex clustering (regression) with direct representative selection (DRS) and direct representative prediction (DRP). Within this framework, the model selected representative samples from the training set (DRS) and formed predictions as weighted sums over these representatives (DRP). Consequently, this DRS/DRP design yielded instance-level interpretability and facilitated the extraction of chemically meaningful insights. In prior studies, these patterns were interpreted by human experts. In the present study, we introduced the LLM as an assistive interpreter and demonstrated that both chemical framing (prompt design) and model size systematically shaped the depth of mechanistic insight. Notably, the LLM is not intended to uncover entirely new mechanisms, but rather to assist human interpretation by providing alternative perspectives, which may help reveal implicit cognitive biases and support more balanced mechanistic reasoning. Specifically, stronger framing and larger models elicited more mechanism-oriented reasoning, whereas weaker framing or smaller models produced concise but more surface-level summaries. Altogether, DRS/DRP enabled a two-layer interpretability framework that linked quantitative attribution from the interpretable ML layer (modified convex regression) with the LLM's linguistic, mechanism-oriented analysis, thereby enabling structured extraction of physicochemical insights from datasets. Within this framework, mechanistic interpretations are systematically structured and accumulated with LLM assistance, providing a pathway toward future knowledge discovery.
Related Concept Videos
Language and Cognition
Improving Translational Accuracy
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...