Related Experiment Video
Updated: Sep 11, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Finding the dark matter: Large language model-based enzyme kinetic data extractor and its validation
Galen Wei1, Xinchun Ran1, Runeem Ai-Abssi1
1Department of Chemistry, Vanderbilt University, Nashville, Tennessee, USA.
Abstract:
Despite the vast number of enzymatic kinetic measurements reported across decades of biochemical literature, the majority of relational enzyme kinetic data-linking amino acid sequence, substrate identity, kinetic parameters, and assay conditions-remains uncollected and inaccessible in structured form. This constitutes a significant portion of the "dark matter" of enzymology. Unlocking these hidden data through automated extraction offers an opportunity to expand enzyme dataset diversity and size, critical for building accurate, generalizable models that drive predictive enzyme engineering. To address this limitation, we built EnzyExtract, a large language model-powered pipeline that automates the extraction, verification, and structuring of enzyme kinetics data from scientific literature. By processing 137,892 full-text publications (PDF/XML), EnzyExtract collected more than 218,095 enzyme-substrate-kinetics entries, including 218,095 kcat and 167,794 Km values. These entries are mapped to enzymes spanning 3569 unique four-digit EC numbers, with a total of 84,464 entries assigned at least a first-digit EC number. EnzyExtract identified 89,544 unique kinetic entries (kcat and Km combined) absent from BRENDA, significantly expanding the known enzymology dataset. The newly curated dataset was compiled into a database named EnzyExtractDB. EnzyExtract demonstrates high accuracy when benchmarked against manually curated datasets and strong consistency with BRENDA-derived data. To create model-ready datasets, enzyme and substrate sequences were aligned to UniProt and PubChem, yielding 92,286 high-confidence, sequence-mapped kinetic entries. To assess the practical utility of our dataset, we retrained several state-of-the-art kcat predictors (including MESI, DLKcat, and TurNuP) using EnzyExtractDB. Across held-out test sets, all models demonstrate improved predictive performance in terms of RMSE, MAE, and R2, highlighting the value of high-quality, large-scale, literature-derived EnzyExtractDB for enhancing predictive modeling of enzyme kinetics. The EnzyExtract source code and the database are openly available at https://github.com/ChemBioHTP/EnzyExtract, and an interactive demo can be accessed via Google Colab at https://colab.research.google.com/drive/1MwKSEZzLPNOseksRshbzkkFoO_cgJhva.
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Determination of Michaelis Constant and Maximum Elimination Rate
These parameters can be estimated by analyzing plasma concentration data post-drug administration. A notable example of this application is phenytoin, a drug with capacity-limited kinetics. It's recommended that phenytoin should be administered at two...
Mechanistic Models: Compartment Models in Individual and Population Analysis
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...

