Related Experiment Video
Updated: Aug 5, 2026

Multiscale Sampling of a Heterogeneous Water/Metal Catalyst Interface using Density Functional Theory and Force-Field Molecular Dynamics
Published on: April 12, 2019
Leveraging Large Language Models for Understanding Fundamental Principles of Catalysis
Shane S Michtavy1, Sinhara M H D Perera1, Marc D Porosoff1
1Department of Chemical and Sustainability Engineering, University of Rochester, Rochester, New York 14627, United States.
Large language models (LLMs) can unify heterogeneous catalysis data, overcoming challenges like small datasets and non-standard representations. This enables better AI-driven discovery by linking diverse chemical information for improved catalyst design.
Area of Science:
- Catalysis
- Artificial Intelligence
- Materials Science
Background:
- Heterogeneous catalysis research faces challenges with small, inconsistent datasets and non-standardized catalyst representations.
- Extracting fundamental knowledge requires integrating performance, characterization, and mechanistic data across scales.
- Current AI approaches struggle with the complexity and heterogeneity of catalysis data.
Purpose of the Study:
- To explore the potential of large language models (LLMs) in advancing heterogeneous catalysis research.
- To identify key opportunities for LLMs in standardizing data and extracting knowledge.
- To discuss the integration of LLMs with existing scientific validation methods.
Main Methods:
- Leveraging language as a unifying representation for diverse catalytic data modalities.
- Focusing on three core LLM applications: text-to-properties, text-to-structure, and text-to-mechanistic models.
- Examining LLM-readiness of data and alignment with scientific principles.
Main Results:
- LLMs can standardize dispersed experimental results, making them accessible for statistical modeling.
- Opportunities exist for LLMs to translate text into catalyst properties, structures, and mechanistic insights.
- Coupling LLM-generated hypotheses with physics-grounded validation yields verifiable and actionable representations.
Conclusions:
- Large language models offer a powerful approach to unify and analyze heterogeneous catalysis data.
- Standardizing data representation through LLMs enhances AI's role in catalyst discovery.
- Integrating LLMs with experimental validation bridges lab-scale findings to industrial applications.
Related Concept Videos
Introduction to Mechanisms of Enzyme Catalysis
Introduction to Mechanisms of Enzyme Catalysis
Catalysis
Catalysis
Heterogeneous Catalysis
Catalytically Perfect Enzymes

