Accelerating Catalysis Understanding via Large Language Model Data Extraction and Shallow Machine Learning Techniques

Brianna R Farris1,2, Kevin C Leonard1,2

  • 1Department of Chemical & Petroleum Engineering, The University of Kansas, 4132 Learned Hall 1530 W 15th St, Lawrence, Kansas 66045, United States.

JACS Au
|November 28, 2025
PubMed
Summary

This study introduces a framework using large language models to create large datasets for catalyst design. Shallow learning models then extract valuable insights from this data, accelerating experimental research in catalysis.

Related Concept Videos

Introduction to Mechanisms of Enzyme Catalysis01:13

Introduction to Mechanisms of Enzyme Catalysis

For many years, scientists thought that enzyme-substrate binding took place in a simple "lock-and-key" fashion. This model stated that the enzyme and substrate fit together perfectly in one instantaneous step. However, current research supports a more refined view scientists call induced fit. The induced-fit model expands upon the lock-and-key model by describing a more dynamic interaction between enzyme and substrate. As the enzyme and substrate come together, their interaction causes...
10.4K
Catalysis02:50

Catalysis

The presence of a catalyst affects the rate of a chemical reaction. A catalyst is a substance that can increase the reaction rate without being consumed during the process. A basic comprehension of a catalysts’ role during chemical reactions can be understood from the concept of reaction mechanisms and energy diagrams.
30.0K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving01:29

Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving

Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
264
Catalytically Perfect Enzymes01:07

Catalytically Perfect Enzymes

The theory of catalytically perfect enzymes was first proposed by W.J. Albery and J. R. Knowles in 1976. These enzymes catalyze biochemical reactions at high-speed. Their catalytic efficiency values range from 108-109 M-1s-1. These enzymes are also called 'diffusion-controlled' as the only rate-limiting step in the catalysis is that of the substrate diffusion into the active site. Examples include triose phosphate isomerase, fumarase, and superoxide dismutase.
 
Most enzymes...
4.9K
Turnover Number and Catalytic Efficiency01:19

Turnover Number and Catalytic Efficiency

The turnover number of an enzyme is the maximum number of substrate molecules it can transform per unit time. Turnover numbers for most enzymes range from 1 to 1000 molecules per second. Catalase has the known highest turnover number, capable of converting up to 2.8×106 molecules of hydrogen peroxide into water and oxygen per second. Lysozyme has the lowest known turnover number of half a molecule per second.
Chymotrypsin is a pancreatic enzyme that breaks down proteins during digestion....
19.9K
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K