Related Experiment Video
Updated: Jun 27, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
682
Advancing plant metabolic research by using large language models to expand databases and extract labeled data
Rachel Knapp1, Braidon Johnson1,2, Lucas Busta1
1Department of Chemistry and Biochemistry University of Minnesota Duluth Duluth Minnesota USA.
Applications in Plant Sciences
|August 6, 2025
Summary
Large language models (LLMs) show promise in plant science by extracting data from literature to build specialized databases. This research demonstrates LLM accuracy in identifying enzyme-product pairs, compound-species associations, and transcribing table images.
Area of Science:
- Plant science
- Metabolomics
- Bioinformatics
Background:
- Scalable data collection in plant science has advanced significantly.
- Machine learning applied to large datasets provides valuable insights in plant metabolic research.
- Large language models (LLMs) offer a new frontier for consolidating scientific literature.
Purpose of the Study:
- To evaluate the efficacy of LLMs in extracting structured data from plant science literature.
- To assess LLM performance in identifying enzyme-product pairs, compound-species associations, and transcribing tabular data.
- To explore the potential of LLMs in creating specialized, comprehensive plant science databases.
Main Methods:
- Testing various prompt engineering techniques and language models for enzyme-product pair identification.
- Applying automated prompt engineering and retrieval-augmented generation for compound-species associations.
- Developing and validating a multimodal LLM pipeline for transcribing table images into machine-readable formats.
Main Results:
- High accuracies (80-90%) achieved for enzyme-product pair identification and table image transcription when models are task-tuned.
- Modest accuracy (50%) for table image transcription.
- Reduced false-negative rates (from 55% to 40%) in compound-species pair identification compared to previous methods.
Conclusions:
- LLMs demonstrate significant potential for automating data extraction and knowledge consolidation in plant science.
- Task-specific tuning of LLMs is crucial for achieving high performance.
- Domain expertise remains vital for effective utilization of LLMs in scientific research.

